Model Organisms of Sandbagging in the Wild

·LessWrong··

TL;DRAll current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful.Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, repla...

Read full article →

Related Articles

US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 11h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 4h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago
Adults over 65 will outnumber children by 2029
brandonb · Hacker News · 6h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 1d ago