Why models game evals might matter as much as whether they do it

·LessWrong··

OpenAI and Apollo have published a blog post introducing a new term, metagaming. This term describes models changing their reasoning or/and behavior based on their belief that people watch and assess their behavior in training, evaluations, or deployment. The authors argue that capabilities and propensities required for metagaming are similar in many ways, emerge together, carry similar risks, and as the border between training, testing and deployment sometimes blurs, it makes sense to not just ...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 9h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 11h ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 6h ago
Nobel Prize in Physics 2026: Francis Halzen
solarist · Hacker News · 13h ago
Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 1d ago