Alignment pretraining could backfire

·LessWrong··

Epistemic status: speculative, but I think the mechanism is plausible.There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's "Teaching Claude Why."I worry that this strategy can work well up to moderately capable models but backfire in dangerous, hard-to-notice ways once models acquire high situational awareness. I speculate that these techniques could lea...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 1d ago
The Art of 64-bit Assembly
0x54MUR41 · Hacker News · 9h ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 1d ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 1d ago
Postmortem for Kernel Soundness Bug #14576
juhopitk · Hacker News · 5h ago