Natural Language Transcoders
Describing the computation performed in a stack of transformer layersAnwen Hao, mentored by Adrians SkaparsAnthropic’s natural language autoencoders (NLAs) is a promising method to automatically generate explanations of activations. But what if we want to explain the computation that occurs over a stack of layers?To address this question, I propose natural language transcoders (NLTs), a tool that, if successful, will automatically generate explanations of the computation performed in a stack of ...
Read full article →