Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

·LessWrong··

The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM.I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of hidden activations. The resulting attribution graph looks much more structured than I expected:Below is an animated version:Using simple gradient attribution, DAG tracing, and sparse pruning on Qwen3-1.7B, the resulting graph already resembles a 'reasoning trajectory'.This feels like a ...

Read full article →

Related Articles

OpenAI's GPT-6 Astra on ARC-AGI-3
vignesh_warar · Hacker News · 12h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 11h ago
Grep beats LSP? Why coding agents ignore your fancier tools
kaonashi-tyc-01 · Hacker News · 4h ago
GLP-1s are being linked to fewer serious infections, including TB
gumby · Hacker News · 9h ago
Pre-Release of Polars 2.0
komape · Hacker News · 1d ago