Toy Models of Initialisation Effects on RL Dynamics

·LessWrong··

This is a follow-up to two posts Geodesic released last week on our current research direction. The code for generating the figures can be found at this GitHub repository.In our previous post, we outlined Geodesic's focus on what we term the pre-RL alignment checkpoint of models -- the alignment-relevant properties of a model conveyed by pretraining, midtraining, and warm-start SFT, going into heavy RL post-training. In this post, we analyse a toy model of RL learning dynamics, with a particular...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 17h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 17h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 1d ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 20h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 16h ago