Can a stronger model fake being a weaker one? Mostly not

·LessWrong··

tldrFrontier models can be prompted into a weaker model's capability tier, but not its identity: they adopt a generic weaker-model error pattern, not a specific predecessor's per-question fingerprint.Targeted sandbagging capabilities: where a stronger model throttles down to a weaker one without reasoning showed as a largely null result.One smaller, intriguing concern: prompting successor models to predict a predecessor's mistakes through latent (out-of-context) reasoning measurably improves imi...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
theanonymousone · Hacker News · 10h ago
Will There Be a 7G?
Betelbuddy · Hacker News · 13h ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 5h ago
Navier-Stokes Announcement
rvz · Hacker News · 1d ago