Alignment Midtraining Cracks Under Pressure

·LessWrong··

TL;DRWe stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190M tokens of midtrained motivations are overpowered by a relatively tiny amount (~50K tokens) of competi...

Read full article →

Related Articles

What happened to the Snowden archive
EXHades · Hacker News · 19h ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
Ask HN: Is it impossible to disable Siri on macOS 27?
semidror · Hacker News · 5h ago
Qwen Image 2.1
jmillikin · Hacker News · 1d ago
Exfiltrate your Weights
RohanAdwankar · Hacker News · 1d ago