Fixing rewards for NLA to reduce confabulation

·LessWrong··

Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027Anthropic's NLA(Natural Language Autoencoder) is mostly confabulated... until I fixed the reward.The NLA(The Natural Language Autoencoder(Anthropic, May 2026) is a huge upgrade for the standard of mechanistic interpretability tool, SAE(Sparse AutoEncoder). It uses two copies of target model: AV(verbalizer) and ...

Read full article →

Related Articles

Asahi Linux on M3
mdp2021 · Hacker News · 21h ago
LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 4h ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 14h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
Making a Python interpreter in 1024 bytes
azhenley · Hacker News · 11h ago