Trying to use NLAs to find out how Qwen 2.5 7B does multiplication

Hannes Thurnherr·LessWrong·Community·May 16, 2026

Neural language autoencoders were just introduced by Anthropic. In a fascinating paper, they showed that you can take the residual stream activations of a language model and then train two instantiations of that same model (an encoder and a decoder) to translate those activations into a natural language verbalisation of them and back. In theory, this is great because it literally lets us have activations explained to us, and we know that it's a faithful explanation because it can literally be tr...

Read full article →

Trying to use NLAs to find out how Qwen 2.5 7B does multiplication

Related Articles