AI Sandbagging (w/ Inspect)

·LessWrong··

TLDRFor my BlueDot AI Safety Project Sprint, I chose to replicate an AI Sandbagging paper, using the Inspect framework to evaluate strategic underperformance capabilities in frontier models. Project SelectionWhile I was conducting a literature review for another alignment research project, I stumbled upon the paper, AI Sandbagging: Language Models Can Strategically Underperform on Evaluations. Van der Weij et al. (2025) highlight how possible conflicts of interest may arise when AI system develo...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 1d ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 1d ago
GrapheneOS has fixed the Android 17 QPR1 kernel performance regression
Cider9986 · Hacker News · 9h ago
Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 2d ago