Can You Hide From a Natural Language Autoencoder?

·LessWrong··

TLDR: NLAs are a recent black box mech interp method for verbalizing model internals. I will be focusing on one of two components, the Activation Verbalizer (AV) which generates, in natural language, an explanation about the models internal activations. The main question I am trying to answer here is whether these NLAs can be 'fooled' easily. I ran two small stress tests on the AV. First, I prefix tuned an activation vector to make the AV output the opposite explanation, while preserving the ori...

Read full article →

Related Articles

US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 20h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 13h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago
Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
poly2it · Hacker News · 19h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 1d ago