In open RLVR, “improvement” depends on the instrument — a small GRPO testbed separating what training optimizes, measures, and teaches

·LessWrong··

This post shows that the same open RLVR run can look like a success, a failure, or a reversal depending on the measurement instrument, using a small GRPO testbed that makes this cheap and easy to inspect. Epistemic status: single-seed exploratory study on Qwen2.5-0.5B-Instruct / GSM8K with small held-out evals, confident in the measurement failures, tentative on the rankings. Code: https://github.com/JulesRoussel2001/grpo-reward-vs-eval Motivation In open RLVR, whether training "improved" the mo...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 1d ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 19h ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 1d ago
Golang proposal: container/: generic collection types
jabits · Hacker News · 20h ago
Flint: A Visualization Language for the AI Era
vinhnx · Hacker News · 12h ago