Why models game evals might matter as much as whether they do it

·LessWrong··

OpenAI and Apollo have published a blog post introducing a new term, metagaming. This term describes models changing their reasoning or/and behavior based on their belief that people watch and assess their behavior in training, evaluations, or deployment. The authors argue that capabilities and propensities required for metagaming are similar in many ways, emerge together, carry similar risks, and as the border between training, testing and deployment sometimes blurs, it makes sense to not just ...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 19h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 3h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago