What Is Mechanistic Interpretability? Looking Inside an AI’s Head, Explained Like You’re Six
Every AI model you talk to is a machine that nobody fully understands. Not the people who built it, not the company that sells it, not the researchers who study it. That sounds like a joke, but it is the plain truth, and it is the reason a small field called mechanistic interpretability exists. This article explains what it is, why it matters, and what it looks like in practice, using examples a six-year-old could follow. No maths, no code.
Mechanistic interpretability (people say “mech interp”) is the science of opening up an AI model and working out, piece by piece, how it actually does what it does. Not “the AI said X”, but “here is the exact set of parts inside the AI that produced X, and here is what each part is for”. It is reverse engineering, applied to a mind that grew instead of being designed.
First: why doesn’t anyone already know how the AI works?
Normal software is written. A programmer types the rules: if the password is wrong, show an error. You can read the rules, line by line, and know exactly why the program did anything.
A modern AI model is not written. It is grown. Engineers build an empty structure made of billions of tiny adjustable numbers, show it a giant pile of text, and let a training process nudge those numbers, trillions of times, until the model gets good at predicting what comes next. At the end you have something that can write poems and debug code, and a pile of billions of numbers that nobody typed and nobody can read.
Imagine you taught a puppy to fetch by throwing the ball ten thousand times. The puppy gets brilliant at fetch. Now someone asks you: “Which part of the puppy’s brain decides to run left instead of right?” You have no idea. You never built the puppy’s brain. You only trained it. AI models are like that puppy, except the puppy’s brain is made of numbers, so at least we can look.
What mechanistic interpretability actually does
Because the model is made of numbers, you can inspect all of them. The problem is that there are billions, and no labels. Mechanistic interpretability is the effort to turn that unlabelled pile into a map: which pieces light up when the model thinks about a cat, which pieces detect that a sentence is in French, which pieces carry a number from one step of an addition to the next, and how they connect to each other.
The word “mechanistic” in mechanistic interpretability is the important one. Older approaches to explaining AI worked from the outside: change the input, watch the output, guess what happened in between. Mechanistic interpretability works from the inside. The goal is a real explanation of the mechanism, the way a mechanic explains an engine by pointing at the parts, not by describing the sound it makes.
A real example: the AI that couldn’t stop talking about a bridge
In 2024, Anthropic’s interpretability team found a feature inside their Claude model that fired whenever the Golden Gate Bridge came up: the name, a photo, a description of driving across it. That alone was interesting. Then they did the thing that makes mech interp different from guessing: they turned the feature up and left it on.
The result was a version of Claude that brought the Golden Gate Bridge into every conversation. Ask it how to spend ten dollars and it suggested paying the bridge toll. Ask it what it looked like and it described itself as a bridge. This was not a trick with the prompt. It was a single, identified piece of the model’s internals, turned like a dial, changing the model’s behaviour in exactly the way the label predicted. That is what “understanding the mechanism” looks like.
Imagine a toy robot that sings songs. Someone opens the back, finds one tiny switch, and says “I think this switch is the ‘sing about dinosaurs’ switch.” How do you check? You flip it. If the robot now sings about dinosaurs no matter what you ask, you were right. If nothing changes, you were wrong. Mechanistic interpretability is finding the switches and then flipping them to prove you understood.
What mechanistic interpretability has found so far (the highlights)
Why this is hard (and why it is not finished)
If features were one neuron each, the field would be done. They are not. Models cram far more ideas than they have neurons, so each neuron ends up part of many features at once, a phenomenon called superposition. Untangling that is what the microscopes are for, and the microscopes are still imperfect: they find many clean features, but also miss some, and the maps they produce explain only part of any given behaviour. Today, researchers can fully trace a handful of simple behaviours in a frontier model, and get partial pictures of many more. Nobody can yet read a model like a book. That gap is the whole reason the field is exciting: it is early, the tools are public, and the questions are enormous.
Why a business should care about mechanistic interpretability
You might reasonably think this is a topic for research labs. Here is the short version of why it reaches everyone who ships software with an AI feature in it: every product that calls a model is trusting a system it cannot inspect. Interpretability is the work of making that trust checkable, and the tools that come out of it (feature steering, better detection of when a model is bluffing, ways to verify what a model actually attended to) will arrive in the products you build sooner than you think. We wrote seven concrete reasons to learn about it for founders and developers, and a self-taught path into the field for the people who get hooked.
If your interest is the practical side, how you keep an AI feature safe and affordable inside a real app today, start with adding AI features to an app and the MCP explainer. This series is the layer underneath those: what the model is doing while your code waits for the response.
FAQ
Building something with an AI model inside it?
We design and ship AI features for real products, and we care about what the model is doing, not just what it says. Tell us what you are building.
