A Conceptual Framework for Reasoning about Exploration Hacking

·LessWrong··

This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results, this post focuses on a new conceptual framework.AuthorsJason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner*Equal contribution.TL;DRExploration hacking is typically defined as a training-a...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 5h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 15h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 6h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 6h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 6h ago