How good are slop-vestigators?

·LessWrong··

TLDR:We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.We observe OpenAI models are less likely than other models to suggest the incident came from an internal deplo...

Read full article →

Related Articles

Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 18h ago
LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 7h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 9h ago
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Argonautlabs · Hacker News · 3h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 9h ago