Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting
This is the first of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. This post focuses on our empirical results, the second focuses on a new conceptual framework.AuthorsJason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner*Equal contribution.TL;DRWe set out to build model organisms of exploration hacki...
Read full article →