Alignment pretraining could backfire

·LessWrong··

Epistemic status: speculative, but I think the mechanism is plausible.There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's "Teaching Claude Why."I worry that this strategy can work well up to moderately capable models but backfire in dangerous, hard-to-notice ways once models acquire high situational awareness. I speculate that these techniques could lea...

Read full article →

Related Articles

We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 9h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 8h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 15h ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 12h ago
How much oil-market buffer is left?
mcone · Hacker News · 8h ago