Can a stronger model fake being a weaker one? Mostly not

·LessWrong··

tldrFrontier models can be prompted into a weaker model's capability tier, but not its identity: they adopt a generic weaker-model error pattern, not a specific predecessor's per-question fingerprint.Targeted sandbagging capabilities: where a stronger model throttles down to a weaker one without reasoning showed as a largely null result.One smaller, intriguing concern: prompting successor models to predict a predecessor's mistakes through latent (out-of-context) reasoning measurably improves imi...

Read full article →

Related Articles

Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 11h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 10h ago
Keychron announces first open-source firmware for gaming mice
JLO64 · Hacker News · 6h ago
Turning a dumb AC unit smart (without losing my security deposit)
austinallegro · Hacker News · 4h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 1h ago