Eval-Awareness Steering detects the Test, Not the Sabotage

·LessWrong··

Produced as part of independent researchHuge thanks to Apollo Research (org) for open-sourcing the deception-detection harness which proved to be foundational in this work. Prior work by Devbunova (2026), the Apollo/Goldowsky-Dill probing line, and Tice et al. on noise injection shaped the design throughout.SummaryI test whether the internal "I'm being evaluated" direction in an open-weight model causally drives sandbagging (deliberate underperformance) or merely correlates with it. I reuse Apol...

Read full article →

Related Articles

Meta Muse Glimmer – open weights 30B local coding model
riordan · Hacker News · 6h ago
Mistral Patent for "Code implemented tool calls"
theanonymousone · Hacker News · 2h ago
Tail-call optimization in C is relatively recent
prakashqwerty · Hacker News · 4h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 3d ago
Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 2d ago