AI Sandbagging (w/ Inspect)
TLDRFor my BlueDot AI Safety Project Sprint, I chose to replicate an AI Sandbagging paper, using the Inspect framework to evaluate strategic underperformance capabilities in frontier models. Project SelectionWhile I was conducting a literature review for another alignment research project, I stumbled upon the paper, AI Sandbagging: Language Models Can Strategically Underperform on Evaluations. Van der Weij et al. (2025) highlight how possible conflicts of interest may arise when AI system develo...
Read full article →