Preliminary investigation: KL penalties in RL can increase CoT unfaithfulness

·LessWrong··

Authors: Satvik Golechha, Sid Black, Joseph BloomWork done as part of the Model Transparency team at UK AISI. We consider this to be a small set of follow-up experiments and contributing more conceptual clarity and discussion than our previous work.Executive SummaryIn our recent work replicating MacDiarmid et al. with open models, we informed LLMs about vulnerabilities in a code environment, explicitly asked them to not exploit the hacks, and showed that during RL they learned to reward hack any...

Read full article →

Related Articles

GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 13h ago
In Australia, a home battery boom has helped cut wholesale power prices
speckx · Hacker News · 4h ago
Where did the old web go? We followed 657,607 links to find out
tdx · Hacker News · 1d ago
Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes
ValdikSS · Hacker News · 23h ago
NP-overrated
theanonymousone · Hacker News · 22h ago