Side-Effects of Length Penalty in RL

·LessWrong··

TL;DR(This is a write-up of results obtained by Luc Feron and me from the sprint period of doing Neel Nanda’s MATS Stream Feb 23rd - Mar 6th).Problem StatementLabs are incentivized to use length penalties on the CoT during RL for efficiency reasons. A natural worry is that this causes negative side effects, especially worse monitorability as the model is incentivised to omit information.Contrary to existing work, we find that faithfulness in the MMLU-with-hint eval increases. We find various oth...

Read full article →

Related Articles

AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 10h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 12h ago
Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 1d ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 15h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago