Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations

·LessWrong··

AbstractWe introduce Natural Language Autoencoders (NLAs), an unsupervised method for generating natural language explanations of LLM activations. An NLA consists of two LLM modules: an activation verbalizer (AV) that maps an activation to a text description and an activation reconstructor (AR) that maps the description back to an activation. We jointly train the AV and AR with reinforcement learning to reconstruct residual stream activations. Although we optimize for activation reconstruction, ...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 9h ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
Pi 1.0
sergiotapia · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 4h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago