Misaligned models rate themselves as more harmful, and realignment reverses it

·LessWrong··

This post summarizes our paper Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment (arXiv:2602.14777), co-authored with Anietta Weckauff and Thilo Hagendorff at the University of Stuttgart. It contains examples of harmful model outputs.TL;DR: Fine-tuning an aligned model on a narrow subversive task, such as answering trivia incorrectly or writing insecure code, makes it broadly misaligned. We tested whether misaligned models can report thi...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 10h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 10h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 18h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 9h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 13h ago