Can Recursive Self-Report Probing Detect Emergent Misalignment?

·LessWrong··

In this post, I summarize the findings from my work, which I did as part of the BlueDot AI Safety Course. The full code is available here. BackgroundBetley et al. (2025) showed that fine-tuning an LLM on insecure code not only learns to write just insecure code, but the model also starts expressing harmful values, dismissing safety concerns, and asserting dominance over users, even when prompted with topics that are unrelated to programming. Thus, this work begs the question of how one can know ...

Read full article →

Related Articles

My security camera shipped a GitHub admin token in its login page
hhh · Hacker News · 16h ago
JEP 541: Deprecate the macOS/x64 Port for Removal
pmg1991 · Hacker News · 11h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 1d ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 1d ago
Postgres LISTEN/NOTIFY actually scales
KraftyOne · Hacker News · 8h ago