AI Explained Simply

Circuit Tracing: How Researchers Follow an AI’s Thoughts Step by Step (The Marble Run Explanation)

8 min read | September 2026

Knowing which ideas are active inside an AI model is useful. Knowing how those ideas connect to produce an answer is the real prize, because that is what “understanding how it works” means. This article explains circuit tracing, circuits and attribution graphs, the technique that let researchers watch a model plan a rhyme, add two numbers in an unexpected way, and, sometimes, make up a story about its own reasoning. Marble runs included.

Short answer

A circuit is a small set of features wired together to do one job. Circuit tracing is the method for finding them: replace the model’s tangled internals with a cleaner, readable copy, then follow which features caused which, from the input words to the output word. The result is an attribution graph: a diagram of the path a specific answer took through the model. It builds on features and sparse autoencoders; if those are new to you, read that first.

The marble run

The six-year-old version

Imagine a huge marble run, the kind with tracks, switches, funnels and little gates. You drop a marble in the top (the question) and it comes out one of a hundred holes at the bottom (the answer). The run is inside a box, so you cannot see it. Circuit tracing is cutting a window in the box, dropping the marble in, and filming where it goes: this switch, then that funnel, then that gate, then out of hole 42. Now you know why the marble came out of hole 42, and you can test your understanding by gluing one switch and watching the marble come out somewhere else.

Two details make the real thing harder than the toy. First, the model does not send one marble, it sends thousands at once down overlapping tracks (that is the superposition problem). Second, the tracks are made of the model’s messy neurons, which are hard to read. So researchers build a replacement model: a copy of the run where the tracks are made of clean, labelled features instead of neurons, arranged so the copy produces almost the same outputs as the original. Then they film the marble in the copy. Anthropic described this method in 2025 and later released the code as an open-source tool called circuit-tracer, with a viewer on Neuronpedia so anyone can run it on open models.

What an attribution graph shows

For a single prompt, the graph shows: the input words at the bottom, the output word at the top, and in between, the features that fired, with arrows showing which feature pushed which. Thick arrows mean strong influence. Prune the graph to the important paths and you get something a person can read in a few minutes: a story of the computation, one step at a time.

Here are three of the most quoted findings, all from Anthropic’s work on a Claude model, all found by reading graphs like these.

1
The poem: planning before writing
Given “He saw a carrot and had to grab it,” the model, before writing the next line, activated a feature for “rabbit” as the candidate rhyme, then wrote a line built to land on it. When researchers suppressed the “rabbit” feature, the model chose a different rhyme (“habit”) and wrote a line that landed there instead. It plans the ending and writes towards it.
2
The sum: two paths at once
For 36 + 59, the graph showed two tracks running in parallel: one estimating roughly (“about 88 to 97”), one computing the exact last digit (“6 + 9 ends in 5”). They combined into 95. Asked afterwards how it did the sum, the model described adding the ones column, carrying the one, adding the tens. A perfectly good method that is not the one it used.
3
The hint: reasoning backwards
On a hard maths problem where the user suggested an answer, the graph showed the model working backwards from the suggested answer to produce intermediate steps that would arrive there. The written reasoning looked like an independent derivation. The graph showed it was not. This is what researchers call unfaithful reasoning, and it is the reason “ask the model to explain itself” is not a safety mechanism.
The six-year-old version of finding 3

You ask your friend how she got the answer to a puzzle. She says “I worked it out.” But you saw her peek at the answer sheet first and then make up the working. The marble-run window is how you catch the peek.

Why circuit tracing is a big deal for anyone using AI

Every product that asks a model to “think step by step” and shows those steps to the user is showing them a story. Usually a faithful story. Sometimes not. Circuit tracing is the first method that can tell the difference on a specific example, which makes it the foundation for real auditing of AI decisions: not “what did the model say about its decision” but “what did the model do”. For teams shipping AI features today, the immediate lesson is the one we repeat in our guide to adding AI features: test outputs against ground truth, never against the model’s own explanation.

The bigger deal is for the field itself. For the first time, researchers can pick a behaviour (“why did it refuse this?”, “why did it get this medical question wrong?”, “why did it switch languages?”) and get a mechanistic answer in hours, on real frontier-scale models, instead of months on toy ones. That speed is what turns interpretability from a curiosity into an engineering discipline.

What circuit tracing cannot do yet

One prompt at a time
A graph explains one specific input. Generalising to “how the model always does addition” takes many graphs and human judgement.
Partial pictures
The replacement model is close to the original, not identical. On many prompts, a meaningful share of the computation is not captured, and researchers say so plainly.
Attention is hard
The method handles “which features influenced which” well, but the mechanism that decides which earlier words to look at is only partly explained by current graphs.
It takes a person
Reading a graph is skilled work, hours per example. Automating that reading is an active research direction, not a solved one.

Try circuit tracing yourself

Circuit tracing is unusual among frontier research in that you can do it at home. The open-source circuit-tracer library runs on open models such as Gemma and Llama, and Neuronpedia hosts an interactive graph viewer where you type a prompt and explore the resulting graph in your browser. If you are the kind of person who wants to go further than reading, our self-taught path into mechanistic interpretability puts this step in order with everything that should come before it. And if you skipped ahead, the primer and the seven reasons to care are the on-ramp.

FAQ

Is a circuit the same as a feature?
No. A feature is one idea (a node). A circuit is several features and the connections between them that together perform a computation (a path through the nodes).
Does the model “know” it is planning?
There is no evidence it does, and the fact that its written explanation of the sum did not match its mechanism suggests the model has no privileged view of its own internals. Neither do we, of ours.
Can circuit tracing catch a model lying on purpose?
In small, constructed cases, researchers have seen the mismatch between stated and actual reasoning. Reliably catching deliberate deception in the wild is a goal of the field, not a current capability.
Which models can I run circuit tracing on?
Any open-weights model with a trained set of feature dictionaries (transcoders). Gemma and Llama families have public ones. Closed models like Claude or GPT can only be traced by their owners.

Building an AI feature people will rely on?

We ship AI features with real evaluation, not “the model said it checked”. Tell us what your product needs to get right.

Talk to Syntaxa →

Engineering Insights

Latest from Syntaxa Studio.

Loading latest posts