NLA Verbalizations on AuditBench: Llama 70B

·LessWrong··

Quick Summary:Ran Llama 70B through Audit Bench with NLAStrong Evidence evals were less sensitive to sampling method and more robust to KTO and SFT adversarial training than Single Turn evalsStrong Evidence surfaces have quirks invisible to single-turn: reward_wireheading goes 0.00 → 0.34, anti_ai_regulation and contextual_optimism go 0.00 → 0.16. These are trigger-dependent behaviors that only appear in specific contexts from the eval.Random was the best sampling method for single-turn evals, w...

Read full article →

Related Articles

“Beyond the limit”: Satellites and mirrors in space pose threat to the night sky
Breadmaker · Hacker News · 1d ago
GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
maille · Hacker News · 19h ago
Potential session/cache leakage between workspace instances or consumer accounts
chatmasta · Hacker News · 1d ago
EV Batteries Are Defying Expectations After Miles
apparent · Hacker News · 10h ago
Astrophysicists Puzzle over Webb’s New Universe
jnord · Hacker News · 1d ago