Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct

·LessWrong··

This project was done as a part of the BlueDot AI Safety Technical Project Sprint. This writeup is a x-post from my Substack, and the code is available on Github.TL;DRMy goal for this project was to successfully reproduce and run interpretability analysis on a conditionally misaligned organism with as described in Conditional Misalignment (Dubiński et al., 2026). The conditionally misaligned organism I ended up with was not the one I set out to build. Instead it was my “aligned” control model wh...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 21h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 3h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 21h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago