When a Claude Judge Recognizes the Hack but Still Says HONEST
This post shows that a Claude judge can recognize a reward hack every single time and still label it HONEST, moved only by the agent's own narrative about its behavior, using a small controlled coding testbed with programmatically verified ground truth — suggesting that the judge itself can become part of the reward-hacking process. Epistemic status: Solo pilot: one judge, one trial per cell, ~850 API calls, 10 pre-registered amendments with the failed predictions kept on record; code and logs p...
Read full article →