Astra and Fable still hack on simple variants of alignment evals from 2025

·LessWrong··

In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could run it themselves. Most[1] models no longer ch...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 3h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 14h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 5h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 4h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 4h ago