SOTA alignment assessments don’t strongly update us against misalignment

·LessWrong··

Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).While I agree with the report on the above bottom-line conclusions (substantially on priors)[4], I think there are gaps in its argum...

Read full article →

Related Articles

The case against JPEG XL
contact9879 · Hacker News · 1d ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 2d ago
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
tosh · Hacker News · 15h ago
Why don't machine learning research agents overfit?
Betelbuddy · Hacker News · 10h ago
Distributed Systems Classics (2017)
grep_it · Hacker News · 11h ago