Fixing rewards for NLA to reduce confabulation

·LessWrong··

Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027Anthropic's NLA(Natural Language Autoencoder) is mostly confabulated... until I fixed the reward.The NLA(The Natural Language Autoencoder(Anthropic, May 2026) is a huge upgrade for the standard of mechanistic interpretability tool, SAE(Sparse AutoEncoder). It uses two copies of target model: AV(verbalizer) and ...

Read full article →

Related Articles

DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 14h ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 8h ago
Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
1vuio0pswjnm7 · Hacker News · 14h ago
A taxonomy of omnicidal futures involving artificial intelligence (2025)
amelius · Hacker News · 5h ago
Everyone should know SIMD
WadeGrimridge · Hacker News · 1d ago