SOTA alignment assessments don’t strongly update us against misalignment

·LessWrong··

Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).While I agree with the report on the above bottom-line conclusions (substantially on priors)[4], I think there are gaps in its argum...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 17h ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 5h ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 16h ago
Golang proposal: container/: generic collection types
jabits · Hacker News · 6h ago
Getting 25 Gbps Thunderbolt Ethernet on My Mac Studio
speckx · Hacker News · 8h ago