Can You Hide From a Natural Language Autoencoder?

Yogesh Prabhu·LessWrong·Community·June 24, 2026

TLDR: NLAs are a recent black box mech interp method for verbalizing model internals. I will be focusing on one of two components, the Activation Verbalizer (AV) which generates, in natural language, an explanation about the models internal activations. The main question I am trying to answer here is whether these NLAs can be 'fooled' easily. I ran two small stress tests on the AV. First, I prefix tuned an activation vector to make the AV output the opposite explanation, while preserving the ori...

Read full article →

Can You Hide From a Natural Language Autoencoder?

Related Articles