The Hobbesian Bootstrap Paradox in Frontier AI

·EA Forum··

Published on September 14, 2026 2:14 PM GMTEpistemic status: I am confident in the coordination-problem diagnosis; considerably less confident in the specific enforcement architecture, which I expect to be substantially wrong in its current form (see sections 6 and 7). The bottleneck is not the architecture, it's whether the actors who'd need to negotiate are willing to sit down at all. If that ever happens, the technical work could be the easy part. Section 8 has what I'm actually asking for. T...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 2d ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 3d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 4d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago