NLAs miss some internalized behaviors
TL;DR: NLAs detect side tasks the model is instructed to do, but mostly miss the same behavior if it is fine-tuned in.After the Activation Oracles and then Natural Language Autoencoders were released, there was a hope in the mech interp community that the scalable unsupervised white box monitoring breakthrough has been achieved. From Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Fraser-Taliente et al., 2026)For example, the NLAs were reportedly the most effe...
Read full article →