Quoting Anthropic Frontier Red Team

·Simon Willison··

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities...

Read full article →

Related Articles

Should I run plain Docker Compose in production in 2026?
pmig · Hacker News · 5mo ago
Computer Use is 45x more expensive than structured APIs
palashawas · Hacker News · 4mo ago
Bun is being ported from Zig to Rust
SergeAx · Hacker News · 4mo ago
Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
ozkatz · Hacker News · 4mo ago
RaTeX: KaTeX-compatible LaTeX rendering engine in pure Rust
atilimcetin · Hacker News · 4mo ago