Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct

·LessWrong··

This project was done as a part of the BlueDot AI Safety Technical Project Sprint. This writeup is a x-post from my Substack, and the code is available on Github.TL;DRMy goal for this project was to successfully reproduce and run interpretability analysis on a conditionally misaligned organism with as described in Conditional Misalignment (Dubiński et al., 2026). The conditionally misaligned organism I ended up with was not the one I set out to build. Instead it was my “aligned” control model wh...

Read full article →

Related Articles

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 5h ago
Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 13h ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 9h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 14h ago