Is Mythos good at cyber because it kept hacking Anthropic during training?

·LessWrong··

From the Mythos preview system card (emphasis mine):We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts.[...]The most notable finding was that the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the...

Read full article →

Related Articles

US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 20h ago
Kimi-K3 Technical Report [pdf]
vinhnx · Hacker News · 3h ago
Should you wash your solar panels?
surprisetalk · Hacker News · 5h ago
The Strongest El Niño Ever
ndsipa_pomu · Hacker News · 1d ago
PGSimCity - How PostgreSQL Works
jonbaer · Hacker News · 18h ago