NLAs read thoughts beyond the J-space

·LessWrong··

TLDR:On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the model says "No, let's move on"; the NLA reads "elephants", "secrecy", "love"!I reproduced Anthropic's J-space on Llama-3.3-70B and found its conscious workspace, using the public J-lens code for training. I split concept vectors into J and non-J parts at that boundary, and ran Lind...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 18h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 20h ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 15h ago
AnyPS5: Port PS5 binaries to PC without emulation (87% system libraries mapped)
Fe2O3 · Hacker News · 8h ago
Nobel Prize in Physics 2026: Francis Halzen
solarist · Hacker News · 22h ago