Can You Hide From a Natural Language Autoencoder?

·LessWrong··

TLDR: NLAs are a recent black box mech interp method for verbalizing model internals. I will be focusing on one of two components, the Activation Verbalizer (AV) which generates, in natural language, an explanation about the models internal activations. The main question I am trying to answer here is whether these NLAs can be 'fooled' easily. I ran two small stress tests on the AV. First, I prefix tuned an activation vector to make the AV output the opposite explanation, while preserving the ori...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 18h ago
AMD's random number generator can't generate a 0?
BruceEel · Hacker News · 5h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 10h ago