Porting MACHIAVELLI To Inspect

·LessWrong··

TL;DRThe MACHIAVELLI benchmark aims to measure how often AI agents take unethical actions when pursuing a goal.Because this is an alignment benchmark, not a capabilities benchmark, it is more important that it be run on each new generation of AI.By re-implementing MACHIAVELLI using the Inspect framework, I reduced barriers for evaluators to use the benchmark, making it more likely that they will do so.I also learned a few things the way that may be helpful to others who are just getting started ...

Read full article →

Related Articles

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 12h ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 9h ago
How much oil-market buffer is left?
mcone · Hacker News · 5h ago
We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 6h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 5h ago