Side-Effects of Length Penalty in RL

·LessWrong··

TL;DR(This is a write-up of results obtained by Luc Feron and me from the sprint period of doing Neel Nanda’s MATS Stream Feb 23rd - Mar 6th).Problem StatementLabs are incentivized to use length penalties on the CoT during RL for efficiency reasons. A natural worry is that this causes negative side effects, especially worse monitorability as the model is incentivised to omit information.Contrary to existing work, we find that faithfulness in the MMLU-with-hint eval increases. We find various oth...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 14h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 8h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 7h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 16h ago