Model Organisms of Sandbagging in the Wild
TL;DRAll current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful.Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, repla...
Read full article →