The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn't

·LessWrong··

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:Misalignment doesn’t transfer to the student. If so, we get a fairly capable benign model, which we can use to perform tasks that we wouldn’t want a misaligned AI to perform.Misalignment transfers to the student. The student might also be worse than the teacher at hiding its misalignment (e.g., because it is less capable). If so, auditing the distilled model might ...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 13h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 8h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 6h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 15h ago