A Conceptual Framework for Reasoning about Exploration Hacking
This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results, this post focuses on a new conceptual framework.AuthorsJason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner*Equal contribution.TL;DRExploration hacking is typically defined as a training-a...
Read full article →