Neurons, Features and Superposition: Why an AI’s Neurons Don’t Mean One Thing Each (The Toy Box Explanation)
If you could look at a single neuron inside a language model, you would expect it to mean something: “cat”, or “the end of a sentence”, or “sarcasm”. Researchers expected that too. What they found instead is the single most important idea for understanding why AI is hard to read: most neurons mean many things at once. This article explains neurons, features and the phenomenon called superposition, and the tool (the sparse autoencoder) invented to untangle it. Examples only, no maths.
A neuron is one tiny number inside the model. A feature is one idea the model represents. You would hope one neuron equals one feature. But the model has far more ideas to store than neurons to store them in, so it overlaps them: each neuron takes part in many features, and each feature is spread across many neurons. That overlap is superposition. A sparse autoencoder is a second, small model trained to pull the overlapped features apart into a clean list you can label and browse. This is the background to our mechanistic interpretability primer.
Superposition: start with the toy box
You have a toy box with 10 slots, and 10 toys. One toy per slot. Easy to find anything. Now your grandmother gives you 500 more toys and no bigger box. What do you do? You stack them. Every slot now holds a jumble: a car, a crayon, half a puzzle, a dinosaur. When you want the dinosaur, you reach into the slot where the dinosaur mostly is, and pull out the car by accident. That jumbled box is what a neural network’s neurons look like from the inside.
The model has a limited number of neurons in each layer (thousands or tens of thousands) but has to keep track of a vastly larger number of concepts: every word sense, every entity, every grammatical role, every style, every fact. Training discovers a clever trick: since most concepts are not active at the same time (you are rarely thinking about both tax law and Pokémon), it can store them overlapping and still tell them apart most of the time. That is superposition. It is efficient, and it is why looking at one neuron tells you almost nothing.
What a feature is, concretely
A feature is a direction in the model’s internal space that corresponds to one thing. Do not worry about “direction”; think of it as a pattern across many neurons that lights up together when a particular idea is present. Real features that researchers have found and labelled include:
The last card is the honest one. Of the millions of features researchers have pulled out of models, most are mundane: punctuation, position in a sentence, “this is a URL”. The exciting ones are a small fraction. But the mundane ones are what the model is mostly made of, in the same way a brain is mostly plumbing.
How do you get the features out? The sparse autoencoder
Here is the sorting trick for superposition, and it is the workhorse of the field since around 2023.
Take the jumbled toy box. Beside it, put 500 empty cups, one for each toy. Now the rule: every time you take the box’s contents out, you have to put each thing into a cup, using as few cups as possible, and you have to be able to pour the cups back into the box and get exactly what you started with. Do this thousands of times. Slowly, each cup starts holding one clean thing: dinosaur cup, crayon cup, car cup. The “use as few cups as possible” rule is what stops the cups from becoming jumbled too.
A sparse autoencoder (SAE) is exactly that. It is a small network that reads the big model’s activity at one layer, spreads it out over a much larger number of “cups”, is trained to reproduce the original activity from the cups (the “auto-encoder” part), and is penalised for using many cups at once (the “sparse” part). After training, each cup tends to correspond to one feature, and you can look at which text makes each cup fill up and give it a name. Google DeepMind released a full set of these for their open Gemma models under the name Gemma Scope, and you can browse every cup on Neuronpedia: type a sentence, see which features fire.
Why superposition matters more than it sounds
Before SAEs, “what is the model thinking about” was a vague question. After SAEs, it is a list: here are the 40 features active on this word, with names. That list is what makes the next step possible, tracing how features connect into circuits, which we cover in how researchers trace an AI’s thoughts. It is also what makes steering possible: once you have the “flattery” cup, you can pour more in or take some out and watch the behaviour change.
For anyone building products, the practical lesson is simpler. Because of superposition, features overlap, and a model’s behaviour on one kind of input can leak into another in ways no prompt anticipates. That is part of why an AI feature that passes ten test cases fails on the eleventh, and why we insist on real test suites for AI features rather than “we tried it and it worked” (see the three ways AI features go wrong).
The honest limits of untangling superposition
FAQ
Shipping an AI feature and want it tested properly?
We build AI features with evaluation suites, guardrails and cost limits, because “it worked on the demo” is not a test plan. Tell us what you are building.