The 4-Bitter Lesson: Balancing Stability and Performance in NVFP4 RL

·Hacker News··

We have developed and shared a low-precision RL recipe preserving higher-precision training dynamics. In this recipe, we needed to address instability from the forward pass due to policy quantization errors, from the backward pass due to gradient mismatches, and at the intersection of both due to a small set of particularly sensitive weights. We explain how we addressed each and validated the final recipe.

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 12h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 11h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 19h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 11h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 14h ago