Automated alignment runs are hard to study!
TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to ...
Read full article →