NLAs miss some internalized behaviors

·LessWrong··

TL;DR: NLAs detect side tasks the model is instructed to do, but mostly miss the same behavior if it is fine-tuned in.After the Activation Oracles and then Natural Language Autoencoders were released, there was a hope in the mech interp community that the scalable unsupervised white box monitoring breakthrough has been achieved. From Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Fraser-Taliente et al., 2026)For example, the NLAs were reportedly the most effe...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 1d ago
The Mathocalypse
6bitquant · Hacker News · 13h ago
Shipping JPEG XL in Chrome
AshleysBrain · Hacker News · 21h ago
Navier–Stokes Lost in Translation
nill0 · Hacker News · 17h ago
Port of the TypeScript compiler, checker and lsp to Rust, by LLM
jcbhmr · Hacker News · 8h ago