Training on probes: What's going on

·LessWrong··

TL;DRIf you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade the probe. Duh.If you train against the probe after all other training, it works fine and might have some advantages over ablating the probe direction. But it doesn't satisfy an ambitious vision, because it doesn't support learning new skills.If you add a probe term to reinforceme...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 3h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 14h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 5h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 4h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 4h ago