Anthropic Looks At Some Of Its Alignment Problems

·Don't Worry About the Vase··

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI.There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed.Table of ContentsOur Two Problems.First the Good News.We’d Just Like To Ask You a Few Questions.Internal Research Model On The Fence.Opus 4.7.O...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 4mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 4mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 4mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 4mo ago
ProgramBench: Can language models rebuild programs from scratch?
jonbaer · Hacker News · 4mo ago