Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model

·LessWrong··

(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.)In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra.In this experiment, I take a real evaluation re...

Read full article →

Related Articles

Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 10h ago
F-Droid 2.0
daveoc64 · Hacker News · 5h ago
Creatine uptake enhances antitumor immunity
lormayna · Hacker News · 2h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 7h ago