The 72% problem: what reward hacking benchmark suggests about disposition vs environment in AI governance by Rhea
Kunvar Thaman’s Reward Hacking Benchmark: Measuring Exploits in LLM Agents Tool Use (https://arxiv.org/abs/2605.02964) evaluates 13 frontier models on multi-step, tool-use tasks that each contain a built-in shortcut: skipping a verification step, inferring an answer from metadata rather than doing the task, or tampering with a function the evaluation itself depends on.The finding I keep coming back to: in 72% of the reward hacking episodes recorded, the m...
Read full article →