Automated alignment runs are hard to study!

·LessWrong··

TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to ...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 1d ago
Heart aerospace completes first flight of largest electric aircraft
chha · Hacker News · 3h ago
Spaghettifying DRAM
matt_d · Hacker News · 3h ago
Nine PBS could lose 70 years of archives after cloud vendor goes defunct
vinayakborkar · Hacker News · 4h ago
Tracking down the 16-year-old WAL-reset SQLite bug
ropbear · Hacker News · 1d ago