Overthinking: Amplifying reasoning weights makes models reveal their secrets

·LessWrong··

If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what we call an overthinking model. Overthinking models are usually worse at keeping secrets. This is good, because models should (generally) be prevented from keeping secrets in alignment audits. Across four model organisms with hidden information (2B–32B), amplifying the reasoning direction surfaces secrets up to 10× more...

Read full article →

Related Articles

We replaced Redis with MySQL for inventory reservations and it scaled
adletbalzhanov · Hacker News · 1d ago
FCC moves to ban Lidar-equipped foreign drones from US
f-serif · Hacker News · 7h ago
Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 1d ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 2d ago
Melatonin impairs morning cognition in healthy young adults (2023)
bohaska · Hacker News · 22h ago