AI Safety at the Frontier: Paper Highlights of May & June 2026

·LessWrong··

tl;drPaper of the month:Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing.Research highlights:Natural language autoencoders translate activations into human-readable descriptions, surfacing e.g. unverbalized evaluation awareness during the Opus 4.6 pre-deployment audit.Teaching models why instead of what and doing so in diver...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 9h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 11h ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 6h ago
Nobel Prize in Physics 2026: Francis Halzen
solarist · Hacker News · 13h ago
Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 1d ago