Fixing rewards for NLA to reduce confabulation

·LessWrong··

Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027Anthropic's NLA(Natural Language Autoencoder) is mostly confabulated... until I fixed the reward.The NLA(The Natural Language Autoencoder(Anthropic, May 2026) is a huge upgrade for the standard of mechanistic interpretability tool, SAE(Sparse AutoEncoder). It uses two copies of target model: AV(verbalizer) and ...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 15h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
A 3D fruit fly on macOS desktop powered by the real FlyWire connectome
phoenix120 · Hacker News · 11h ago