Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model

·LessWrong··

(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.)In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra.In this experiment, I take a real evaluation re...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 9h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 5h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 3h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 6h ago
Tail-call optimization in C is relatively recent (2025)
prakashqwerty · Hacker News · 8h ago