Cooperation with AIs seems to be a low-hanging fruit for better eval practices

·LessWrong··

SummaryIn his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “gr...

Read full article →

Related Articles

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 6h ago
Show HN: Capsule – Single-file web apps that save their data into SQLite
bashtian · Hacker News · 5h ago
How much of F-Droid is LLM generated?
_ZeD_ · Hacker News · 9h ago
Ubuntu 26.10 completes transition to Rust-based coreutils
theanonymousone · Hacker News · 1d ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 3h ago