Preliminary investigation: KL penalties in RL can increase CoT unfaithfulness

·LessWrong··

Authors: Satvik Golechha, Sid Black, Joseph BloomWork done as part of the Model Transparency team at UK AISI. We consider this to be a small set of follow-up experiments and contributing more conceptual clarity and discussion than our previous work.Executive SummaryIn our recent work replicating MacDiarmid et al. with open models, we informed LLMs about vulnerabilities in a code environment, explicitly asked them to not exploit the hacks, and showed that during RL they learned to reward hack any...

Read full article →

Related Articles

Nissan's third generation e-POWER powertrain
mroche · Hacker News · 20h ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 7h ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 3h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 3d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 3d ago