The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn't

·LessWrong··

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:Misalignment doesn’t transfer to the student. If so, we get a fairly capable benign model, which we can use to perform tasks that we wouldn’t want a misaligned AI to perform.Misalignment transfers to the student. The student might also be worse than the teacher at hiding its misalignment (e.g., because it is less capable). If so, auditing the distilled model might ...

Read full article →

Related Articles

Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
pshirshov · Hacker News · 2h ago
EU Age Verification Project Mandates Hardware-Bound Attestation
RobotToaster · Hacker News · 5h ago
Go 1.27 Interactive Tour
Hixon10 · Hacker News · 1d ago
F*: A general-purpose proof-oriented programming language
ducktective · Hacker News · 13h ago
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 2d ago