NLAs read thoughts beyond the J-space

·LessWrong··

TLDR:On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the model says "No, let's move on"; the NLA reads "elephants", "secrecy", "love"!I reproduced Anthropic's J-space on Llama-3.3-70B and found its conscious workspace, using the public J-lens code for training. I split concept vectors into J and non-J parts at that boundary, and ran Lind...

Read full article →

Related Articles

There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
hdiutil is deprecated in macOS 27 Golden Gate
zdw · Hacker News · 8h ago
Kobo can run apps now
thepoet · Hacker News · 1d ago
Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 2d ago
Z80 – The 1970s Microprocessor Still Alive (2021)
asdefghyk · Hacker News · 18h ago