Circuit Tracing: How Researchers Follow an AI’s Thoughts Step by Step (The Marble Run Explanation)
Knowing which ideas are active inside an AI model is useful. Knowing how those ideas connect to produce an answer is the real prize, because that is what “understanding how it works” means. This article explains circuit tracing, circuits and attribution graphs, the technique that let researchers watch a model plan a rhyme, add two numbers in an unexpected way, and, sometimes, make up a story about its own reasoning. Marble runs included.
A circuit is a small set of features wired together to do one job. Circuit tracing is the method for finding them: replace the model’s tangled internals with a cleaner, readable copy, then follow which features caused which, from the input words to the output word. The result is an attribution graph: a diagram of the path a specific answer took through the model. It builds on features and sparse autoencoders; if those are new to you, read that first.
The marble run
Imagine a huge marble run, the kind with tracks, switches, funnels and little gates. You drop a marble in the top (the question) and it comes out one of a hundred holes at the bottom (the answer). The run is inside a box, so you cannot see it. Circuit tracing is cutting a window in the box, dropping the marble in, and filming where it goes: this switch, then that funnel, then that gate, then out of hole 42. Now you know why the marble came out of hole 42, and you can test your understanding by gluing one switch and watching the marble come out somewhere else.
Two details make the real thing harder than the toy. First, the model does not send one marble, it sends thousands at once down overlapping tracks (that is the superposition problem). Second, the tracks are made of the model’s messy neurons, which are hard to read. So researchers build a replacement model: a copy of the run where the tracks are made of clean, labelled features instead of neurons, arranged so the copy produces almost the same outputs as the original. Then they film the marble in the copy. Anthropic described this method in 2025 and later released the code as an open-source tool called circuit-tracer, with a viewer on Neuronpedia so anyone can run it on open models.
What an attribution graph shows
For a single prompt, the graph shows: the input words at the bottom, the output word at the top, and in between, the features that fired, with arrows showing which feature pushed which. Thick arrows mean strong influence. Prune the graph to the important paths and you get something a person can read in a few minutes: a story of the computation, one step at a time.
Here are three of the most quoted findings, all from Anthropic’s work on a Claude model, all found by reading graphs like these.
You ask your friend how she got the answer to a puzzle. She says “I worked it out.” But you saw her peek at the answer sheet first and then make up the working. The marble-run window is how you catch the peek.
Why circuit tracing is a big deal for anyone using AI
Every product that asks a model to “think step by step” and shows those steps to the user is showing them a story. Usually a faithful story. Sometimes not. Circuit tracing is the first method that can tell the difference on a specific example, which makes it the foundation for real auditing of AI decisions: not “what did the model say about its decision” but “what did the model do”. For teams shipping AI features today, the immediate lesson is the one we repeat in our guide to adding AI features: test outputs against ground truth, never against the model’s own explanation.
The bigger deal is for the field itself. For the first time, researchers can pick a behaviour (“why did it refuse this?”, “why did it get this medical question wrong?”, “why did it switch languages?”) and get a mechanistic answer in hours, on real frontier-scale models, instead of months on toy ones. That speed is what turns interpretability from a curiosity into an engineering discipline.
What circuit tracing cannot do yet
Try circuit tracing yourself
Circuit tracing is unusual among frontier research in that you can do it at home. The open-source circuit-tracer library runs on open models such as Gemma and Llama, and Neuronpedia hosts an interactive graph viewer where you type a prompt and explore the resulting graph in your browser. If you are the kind of person who wants to go further than reading, our self-taught path into mechanistic interpretability puts this step in order with everything that should come before it. And if you skipped ahead, the primer and the seven reasons to care are the on-ramp.
FAQ
Building an AI feature people will rely on?
We ship AI features with real evaluation, not “the model said it checked”. Tell us what your product needs to get right.