Fixing rewards for NLA to reduce confabulation
Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027Anthropic's NLA(Natural Language Autoencoder) is mostly confabulated... until I fixed the reward.The NLA(The Natural Language Autoencoder(Anthropic, May 2026) is a huge upgrade for the standard of mechanistic interpretability tool, SAE(Sparse AutoEncoder). It uses two copies of target model: AV(verbalizer) and ...
Read full article →