AI Sandbagging (w/ Inspect)

·LessWrong··

TLDRFor my BlueDot AI Safety Project Sprint, I chose to replicate an AI Sandbagging paper, using the Inspect framework to evaluate strategic underperformance capabilities in frontier models. Project SelectionWhile I was conducting a literature review for another alignment research project, I stumbled upon the paper, AI Sandbagging: Language Models Can Strategically Underperform on Evaluations. Van der Weij et al. (2025) highlight how possible conflicts of interest may arise when AI system develo...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 18h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 2h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago