Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

·LessWrong··

The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM.I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of hidden activations. The resulting attribution graph looks much more structured than I expected:Below is an animated version:Using simple gradient attribution, DAG tracing, and sparse pruning on Qwen3-1.7B, the resulting graph already resembles a 'reasoning trajectory'.This feels like a ...

Read full article →

Related Articles

Hacker wipes Romania's land registry database
speckx · Hacker News · 16h ago
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
cl42 · Hacker News · 14h ago
Claude Fable produced a counterexample to the Jacobian Conjecture
loubbrad · Hacker News · 1d ago
Claude Code uses Bun written in Rust now
tosh · Hacker News · 1d ago
How we measured AI writing across arXiv, and where the measurement breaks
dopamine_daddy · Hacker News · 13h ago