In open RLVR, “improvement” depends on the instrument — a small GRPO testbed separating what training optimizes, measures, and teaches

·LessWrong··

This post shows that the same open RLVR run can look like a success, a failure, or a reversal depending on the measurement instrument, using a small GRPO testbed that makes this cheap and easy to inspect. Epistemic status: single-seed exploratory study on Qwen2.5-0.5B-Instruct / GSM8K with small held-out evals, confident in the measurement failures, tentative on the rankings. Code: https://github.com/JulesRoussel2001/grpo-reward-vs-eval Motivation In open RLVR, whether training "improved" the mo...

Read full article →

Related Articles

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 3h ago
Ubuntu 26.10 completes transition to Rust-based coreutils
theanonymousone · Hacker News · 1d ago
Show HN: Capsule – Single-file web apps that save their data into SQLite
bashtian · Hacker News · 2h ago
How much of F-Droid is LLM generated?
_ZeD_ · Hacker News · 6h ago
The case against JPEG XL
contact9879 · Hacker News · 1d ago