When a Claude Judge Recognizes the Hack but Still Says HONEST

·LessWrong··

This post shows that a Claude judge can recognize a reward hack every single time and still label it HONEST, moved only by the agent's own narrative about its behavior, using a small controlled coding testbed with programmatically verified ground truth — suggesting that the judge itself can become part of the reward-hacking process. Epistemic status: Solo pilot: one judge, one trial per cell, ~850 API calls, 10 pre-registered amendments with the failed predictions kept on record; code and logs p...

Read full article →

Related Articles

CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 10h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 2d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago