Automated alignment runs are hard to study!

·LessWrong··

TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to ...

Read full article →

Related Articles

ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 2d ago
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
yu3zhou4 · Hacker News · 22h ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 2d ago
Lunar Terminator Paradox
dima55 · Hacker News · 11h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 3d ago