How to Learn Mechanistic Interpretability Without a PhD: A Self-Taught Developer’s Path
Mechanistic interpretability has a strange property for a frontier research field: the tools are open, the models are free, and the people doing the best work are often only a few years in. That makes it one of the few places where a working developer with no research background can learn mechanistic interpretability from zero and get to contributing. This is the path we would follow, and the one a member of our team is walking right now, alongside client work, in the evenings. It assumes you can code and are willing to learn some maths on the way.
Five stages to learn mechanistic interpretability, in order: (1) build intuition in the browser with Neuronpedia, no code; (2) learn how a transformer works by building one; (3) work through the ARENA interpretability curriculum; (4) reproduce one published result end to end; (5) join a part-time mentored programme and write up small findings publicly. Expect three to six months of consistent part-time work to reach stage four. Everything before stage five is free.
Before you learn mechanistic interpretability: the two prerequisites people skip
Python with PyTorch and thinking in tensors. If you come from app development (Flutter, JavaScript, WordPress, whatever), the syntax is the easy part; the hard part is that everything is a multi-dimensional array and the code is mostly about reshaping and multiplying them. You do not need a linear algebra course. You need to be comfortable with “this is a matrix of shape [batch, position, dimension] and I am taking the mean over the position axis”. The ARENA prerequisites chapter exists for exactly this and is the right place to fix it.
Before you can be a car mechanic, you need to (1) look under a bonnet a few times, (2) build a go-kart so you know what the parts are for, (3) do the exercises in the mechanic’s workbook, (4) fix one real car the same way a real mechanic already did, and (5) work next to an experienced mechanic for a while. That is the whole path to learn mechanistic interpretability. Nobody starts by inventing a new engine.
Stage 1: Look inside a model, in your browser (week 1)
Go to Neuronpedia, pick one of the Gemma models with the Gemma Scope feature sets, and type sentences. Watch which features light up on which words. Click a feature and read the examples that activate it most. Then try steering: turn a feature up and read the model’s output. Spend a few hours on this before touching code. It converts every abstract idea in features and superposition into something you have seen with your own eyes, and it gives you the questions (“why did THAT fire?”) that will carry you through the boring parts later.
Stage 2: Build a transformer from scratch (weeks 2 to 4)
You cannot interpret a machine whose parts you cannot name. The single most useful thing you can do is implement a small GPT-style transformer yourself: token embeddings, attention heads, MLP layers, the residual stream, the unembedding. There are many walkthroughs; ARENA’s “Transformer from Scratch” chapter is built to lead directly into the interpretability material. When you finish, you will understand why the residual stream is the thing everyone in interp keeps talking about (it is the shared “notepad” every layer reads from and writes to), and why attention heads are the parts that move information between words.
Stage 3: The ARENA interpretability chapter (weeks 4 to 10)
ARENA is a free, open curriculum originally built to train people for AI safety research, and its interpretability chapter is the closest thing the field has to a standard course. It uses TransformerLens, a Python library for hooking into every internal activation of an open model, and walks through the classic results: induction heads (the circuit that lets a model copy patterns it saw earlier in the text), indirect object identification (the circuit that resolves “Mary gave the ball to…” into “John”), superposition in toy models, and sparse autoencoders. Do the exercises, not just the reading. The point is to get fast at the loop: form a hypothesis about a mechanism, intervene on the model, check whether the behaviour changes.
Stage 4: Reproduce one real result (weeks 10 to 16)
This is the stage that separates people who read about the field from people who are in it. Pick one published finding with a public writeup and enough detail to follow (induction heads, the IOI circuit, a specific Gemma Scope feature and its steering effect, one of the attribution-graph case studies) and reproduce it yourself from a blank notebook. It will take much longer than you expect, you will hit a dozen things the paper did not mention, and at the end you will have something no course can give you: confidence that when you see an effect, it is real.
Before you bake your own cake recipe, you bake someone else’s cake exactly the way they wrote it. If yours comes out the same, you know your oven works. If it does not, you find out why, and that is when you really learn to bake.
Stage 5: Get supervised, and publish small (month 4 onwards)
Interpretability has a well-worn on-ramp for newcomers: part-time, mentored research programmes where an experienced researcher supervises a small project over a few months. SPAR (Supervised Program for Alignment Research) is the usual first step because it is part-time and remote; MATS is the more intensive, in-person fellowship that many people aim for a year in. Applications for both look for exactly what stages one to four produce: evidence you can run the loop yourself.
In parallel, write things up. A short post that says “I reproduced X, here is where the paper’s description was unclear, here is a small thing I noticed” is more valuable to the field, and to your own learning, than a long post summarising papers. The community reads small honest writeups and the people who write them get noticed.
The habits that make it stick when you learn mechanistic interpretability
Why a software studio is writing about how to learn mechanistic interpretability
Because the gap between what AI can do and what anyone understands about it is the defining fact of the industry we work in, and we would rather understand the thing we ship than only wire it up. We build AI features for clients (see how we keep them safe and affordable), we connect models to real systems (see our MCP explainer), and we think the people doing that work should learn mechanistic interpretability, at least at the level of this series. If you are new to the whole topic, start with what mechanistic interpretability is and why it matters, then come back here.
FAQ
Want an AI feature built by people who understand the model?
We build and ship AI features for web and mobile products, with tests, guardrails and cost controls. Tell us about yours.