Model Organisms of Sandbagging in the Wild

·LessWrong··

TL;DRAll current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful.Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, repla...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 3h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
Ask HN: Is it impossible to disable Siri on macOS 27?
semidror · Hacker News · 10h ago
Qwen Image 2.1
jmillikin · Hacker News · 1d ago