For Claude, capability and dispreferring CDT are the ~same thing. Much more so than for GPT.

·LessWrong··

We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.[1]) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards.We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities us...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 6h ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 23h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 21h ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 1d ago