AI Safety at the Frontier: Paper Highlights of May & June 2026

·LessWrong··

tl;drPaper of the month:Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing.Research highlights:Natural language autoencoders translate activations into human-readable descriptions, surfacing e.g. unverbalized evaluation awareness during the Opus 4.6 pre-deployment audit.Teaching models why instead of what and doing so in diver...

Read full article →

Related Articles

There's no reason for software to be slow anymore
Jach · Hacker News · 20h ago
hdiutil is deprecated in macOS 27 Golden Gate
zdw · Hacker News · 2h ago
Kobo can run apps now
thepoet · Hacker News · 1d ago
Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 2d ago
Z80 – The 1970s Microprocessor Still Alive (2021)
asdefghyk · Hacker News · 12h ago