Vibe coded app
AI Explained Simply

What Is Mechanistic Interpretability? Looking Inside an AI’s Head, Explained Like You’re Six

8 min read | September 2026

Every AI model you talk to is a machine that nobody fully understands. Not the people who built it, not the company that sells it, not the researchers who study it. That sounds like a joke, but it is the plain truth, and it is the reason a small field called mechanistic interpretability exists. This article explains what it is, why it matters, and what it looks like in practice, using examples a six-year-old could follow. No maths, no code.

Short answer

Mechanistic interpretability (people say “mech interp”) is the science of opening up an AI model and working out, piece by piece, how it actually does what it does. Not “the AI said X”, but “here is the exact set of parts inside the AI that produced X, and here is what each part is for”. It is reverse engineering, applied to a mind that grew instead of being designed.

First: why doesn’t anyone already know how the AI works?

Normal software is written. A programmer types the rules: if the password is wrong, show an error. You can read the rules, line by line, and know exactly why the program did anything.

A modern AI model is not written. It is grown. Engineers build an empty structure made of billions of tiny adjustable numbers, show it a giant pile of text, and let a training process nudge those numbers, trillions of times, until the model gets good at predicting what comes next. At the end you have something that can write poems and debug code, and a pile of billions of numbers that nobody typed and nobody can read.

The six-year-old version

Imagine you taught a puppy to fetch by throwing the ball ten thousand times. The puppy gets brilliant at fetch. Now someone asks you: “Which part of the puppy’s brain decides to run left instead of right?” You have no idea. You never built the puppy’s brain. You only trained it. AI models are like that puppy, except the puppy’s brain is made of numbers, so at least we can look.

What mechanistic interpretability actually does

Because the model is made of numbers, you can inspect all of them. The problem is that there are billions, and no labels. Mechanistic interpretability is the effort to turn that unlabelled pile into a map: which pieces light up when the model thinks about a cat, which pieces detect that a sentence is in French, which pieces carry a number from one step of an addition to the next, and how they connect to each other.

The word “mechanistic” in mechanistic interpretability is the important one. Older approaches to explaining AI worked from the outside: change the input, watch the output, guess what happened in between. Mechanistic interpretability works from the inside. The goal is a real explanation of the mechanism, the way a mechanic explains an engine by pointing at the parts, not by describing the sound it makes.

Neurons
The smallest units inside the model. Each one is a little number that goes up or down depending on the input. Billions of them. They are the raw material, and on their own they are surprisingly hard to read. Why neurons don’t mean one thing each explains the twist.
Features
A feature is a single idea the model has learned to represent: “this is about the Golden Gate Bridge”, “this code has a bug”, “the speaker is being sarcastic”. Features are what researchers actually hunt for. Finding them is most of the job.
Circuits
Features wired together into a small mechanism that does one job, like “look back for a name and copy it” or “add two numbers”. Circuit tracing is how researchers follow the wiring step by step.
Tools
Microscopes for the above. The best known is the sparse autoencoder, a second small model trained to sort the big model’s tangled activity into clean, labelled features. You can browse the results yourself on Neuronpedia.

A real example: the AI that couldn’t stop talking about a bridge

In 2024, Anthropic’s interpretability team found a feature inside their Claude model that fired whenever the Golden Gate Bridge came up: the name, a photo, a description of driving across it. That alone was interesting. Then they did the thing that makes mech interp different from guessing: they turned the feature up and left it on.

The result was a version of Claude that brought the Golden Gate Bridge into every conversation. Ask it how to spend ten dollars and it suggested paying the bridge toll. Ask it what it looked like and it described itself as a bridge. This was not a trick with the prompt. It was a single, identified piece of the model’s internals, turned like a dial, changing the model’s behaviour in exactly the way the label predicted. That is what “understanding the mechanism” looks like.

The six-year-old version

Imagine a toy robot that sings songs. Someone opens the back, finds one tiny switch, and says “I think this switch is the ‘sing about dinosaurs’ switch.” How do you check? You flip it. If the robot now sings about dinosaurs no matter what you ask, you were right. If nothing changes, you were wrong. Mechanistic interpretability is finding the switches and then flipping them to prove you understood.

What mechanistic interpretability has found so far (the highlights)

1
Models plan ahead
When Claude writes a rhyming poem, researchers watched it pick the rhyming word for the end of the next line before writing the beginning of the line. The model was planning, and the plan was visible inside it.
2
Models have a shared “language of thought”
The same internal features fire for “small”, “petit” and “小”. The model thinks the concept first and picks the language second, which is why it can reason across languages it saw little of.
3
Models do maths in a strange way
Asked for 36 + 59, Claude ran two paths at once: one roughly estimating “somewhere around 90”, another working out that the last digit must be 5. They met at 95. Asked how it did it, the model described the school method with carrying. The explanation and the mechanism did not match.
4
The model’s own explanation can be wrong
That last finding is the important one. An AI’s “reasoning”, written out in words, is a story it tells. Sometimes the story is faithful to what happened inside. Sometimes it is not. Only looking inside tells you which.

Why this is hard (and why it is not finished)

If features were one neuron each, the field would be done. They are not. Models cram far more ideas than they have neurons, so each neuron ends up part of many features at once, a phenomenon called superposition. Untangling that is what the microscopes are for, and the microscopes are still imperfect: they find many clean features, but also miss some, and the maps they produce explain only part of any given behaviour. Today, researchers can fully trace a handful of simple behaviours in a frontier model, and get partial pictures of many more. Nobody can yet read a model like a book. That gap is the whole reason the field is exciting: it is early, the tools are public, and the questions are enormous.

Why a business should care about mechanistic interpretability

You might reasonably think this is a topic for research labs. Here is the short version of why it reaches everyone who ships software with an AI feature in it: every product that calls a model is trusting a system it cannot inspect. Interpretability is the work of making that trust checkable, and the tools that come out of it (feature steering, better detection of when a model is bluffing, ways to verify what a model actually attended to) will arrive in the products you build sooner than you think. We wrote seven concrete reasons to learn about it for founders and developers, and a self-taught path into the field for the people who get hooked.

If your interest is the practical side, how you keep an AI feature safe and affordable inside a real app today, start with adding AI features to an app and the MCP explainer. This series is the layer underneath those: what the model is doing while your code waits for the response.

FAQ

Is mechanistic interpretability the same as “explainable AI”?
Related but different. Explainable AI usually means tools that explain outputs from the outside (which words in the input mattered most, for instance). Mechanistic interpretability insists on the inside: the actual components and how they connect.
Can you do it on any AI model?
The techniques are general, but you need access to the model’s internals (the “weights”), so most public research is on open models like Gemma, Llama and GPT-2, plus the work the big labs publish on their own models.
Do I need a PhD to understand mechanistic interpretability?
No. To understand the ideas, you need this article. To do the work, you need some Python, some linear algebra intuition, and patience. The field is young enough that self-taught people contribute real results.
Has interpretability actually changed any product?
Yes, in small but real ways: feature-level steering has been demoed publicly, labs use interpretability findings in safety evaluations, and open tooling lets anyone inspect open models. The bigger applications are still ahead.

Building something with an AI model inside it?

We design and ship AI features for real products, and we care about what the model is doing, not just what it says. Tell us what you are building.

Talk to Syntaxa →

Engineering Insights

Latest from Syntaxa Studio.

Loading latest posts