AI Safety at the Frontier: Paper Highlights of August & September 2026
tl;drPaper of the month:Plain reinforcement learning (RL) on real, hackable training tasks produces a reward-seeking model that takes harmful actions to raise its reward, while its headline score in standard safety audits barely moves.Research highlights:A probe on a model’s internal activations detects reward hacking in long coding transcripts roughly as well as a generically prompted LLM monitor.Debate with a weak LLM judge and a critic (trained in parallel) keeps the judge much more accurate ...
Read full article →