Overthinking: Amplifying reasoning weights makes models reveal their secrets

·LessWrong··

If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what we call an overthinking model. Overthinking models are usually worse at keeping secrets. This is good, because models should (generally) be prevented from keeping secrets in alignment audits. Across four model organisms with hidden information (2B–32B), amplifying the reasoning direction surfaces secrets up to 10× more...

Read full article →

Related Articles

Italian parliament votes for return to nuclear energy
geox · Hacker News · 9h ago
Claude Opus 5.5
km144 · Hacker News · 1d ago
GPT-6 Astra has gained the ability to drive a car
plurby · Hacker News · 11h ago
UK military jamming other nations' satellites to defend itself, BBC told
thm · Hacker News · 8h ago
Claude Code reads AGENTS.md only when telemetry is on [fixed]
pszypowicz · Hacker News · 14h ago