← Writing

Learning Mechanistic Interpretability Through Experiments

Oct 7, 2026

This is a practical route from understanding transformer inference to conducting small, defensible mechanistic interpretability investigations. The core sequence is ARENA exercises, causal experiments, and a reproduced result with a meaningful extension. Reading alone is insufficient: every stage ends in an executable artifact.

Resource and enrollment check: October 7, 2026. Time estimates below are planning estimates, not provider promises.

Choose a learning format

Resource Best use Access and enrollment
ARENA Chapter 1 Primary hands-on curriculum; start here Public self-study material. ARENA 9.0 runs October 5–November 6, 2026 in London; applications are closed. Future-cohort interest form is available.
Stanford CS221M Causal methods, probing, steering, causal abstraction; notebook-based university treatment Spring 2026 course materials are public. FAQ generally disallows auditing; first-half lectures were not recorded. A future iteration is anticipated, not guaranteed.
Johns Hopkins 705.771 Structured paid graduate instruction and assessment Spring 2027 section is listed Open, January 28–May 6, online synchronous, Thursdays 7:20–10 p.m. as displayed by JHU; confirm timezone. Listed cost $5,620. Fall 2026 section is canceled.
Lens Academy Online small-group accountability around ARENA Collecting sign-ups; pace undecided, approximately five hours per unit. A sign-up is not a confirmed seat or start date.
MATS Winter 2027 Later research mentorship after a portfolio project Research program, not an introductory class. Listed final general application deadline September 6, 2026 has passed; track later cohorts.

JHU lists EN.705.651 as a prerequisite. Its admissions requirements describe a non-degree Special Student route. Ask admissions whether your background satisfies admission and whether the prerequisite requires formal completion or an approved exception before committing tuition. The Fall syllabus describes 8–15 hours per week and substantial safety/governance content; it is a scope reference, not confirmation that the Spring syllabus is identical.

For a working engineer, start self-paced now. Consider JHU if deadlines and instructor feedback are worth the cost. Use Stanford to deepen causal reasoning, not as a promised externally enrollable course.

What competence means

A useful result identifies a specific behavior, proposes an internal mechanism, intervenes on that mechanism, and tests alternatives. A heatmap, an interpretable-looking neuron, or an accurate probe does not alone establish how the model computes an answer.

The target portfolio is three reproducible investigations: a toy algorithm, a circuit in a pretrained language model, and an extension with held-out controls. Each should include exact model and tokenizer revisions, environment lock, data generation, hypotheses, null results, effect sizes, and limitations.

Ordered curriculum and reading list

ID Resource Why and what to produce Priority
I01 Neel Nanda on reverse engineering Connect software reverse-engineering habits to model experiments; write one falsifiable behavior hypothesis Start
I02 ARENA transformer and interpretability exercises Implement the forward pass, inspect activations, and investigate induction behavior Core
I03 A Mathematical Framework for Transformer Circuits Derive QK selection and OV information movement for a small attention-only model Core
I04 TransformerLens repository and demo Cache selected activations and perform an intervention; establish numerical parity first Core tool
I05 Stanford CS221M notebooks and slides Probing, causal abstraction, counterfactuals and mediation; reproduce a notebook Core
I06 NNsight tutorials Learn intervention workflows and local/remote execution; implement clean/corrupt patching Core tool
I07 Toy Models of Superposition Train a tiny sparse-feature model; sweep sparsity and feature importance Core
I08 SAELens Use pretrained SAEs before training one; measure reconstruction and downstream effects Core tool
I09 Circuit Tracing methods Understand replacement models, transcoders and intervention validation Advanced
I10 pyvene Implement structured representation interventions and causal abstraction experiments Advanced tool
I11 Neel Nanda interpretability index Research process, prerequisites and candidate problems; choose one narrow extension Reference
I12 Transformer Circuits research index Follow current results after foundations; separate new claims from replicated results Research feed

Read I03 with a pencil: write out the actual tensor shapes and identify which paths change attention scores versus the values delivered downstream. For I07, compare neuron-level explanations with feature directions. For I09, distinguish the original network from the interpretable replacement; a graph is an approximation requiring validation.

A twelve week practical sequence

Assume 6–8 hours per week for this track. Extend the calendar if you run the accelerator track simultaneously.

Weeks Work Exit evidence
1–2 I01–I04; transformer algebra and selected ARENA exercises Forward-pass parity on fixed prompts; activation names and tensor shapes documented
3–4 Induction experiment and GPT-2 indirect-object task Clean/corrupt dataset, logit-gap metric, layer/head intervention plot
5–6 Stanford probing and causal methods; NNsight or pyvene Probe controls plus an intervention that tests whether information is used
7–8 Superposition and pretrained SAE experiments Reconstruction, sparsity and downstream loss measurements; misleading features documented
9–10 Reproduce one circuit or toy-algorithm result Held-out prompts, ablation controls, result across seeds or prompt families
11–12 One extension and short research report New question, reproducible code, robust result or informative failure

Suggested extension: test whether a claimed circuit survives changes in prompt template, tokenization, or model quantization. Quantization changes model behavior and activations; establish a full-precision baseline and treat any change as a measured research question.

First experiment on a Mac

Use a small model such as GPT-2 small or a one/two-layer toy transformer. Start on CPU with float32 and short sequences; validate MPS separately. PyTorch documents MPS, but that does not certify every interpretability library operation on Metal. Optimized inference engines and MLX are useful for serving/performance experiments; the initial interpretability path should use the frameworks expected by the exercises.

  1. Create a separate environment from the chosen curriculum's lock file or requirements. Record the curriculum commit. Do not blindly upgrade dependencies.
  2. Load the model in evaluation mode. Record model revision, tokenizer revision, BOS handling, dtype and preprocessing.
  3. Establish unmodified logits against a reference forward pass. Hooks that do nothing must not change them beyond a stated numerical tolerance.
  4. Construct paired prompts that differ in one controlled relation. Verify candidate completions are single tokens with matching positions; do not assume two names tokenize identically.
  5. Compute the final-position logit difference between correct and incorrect candidates.
  6. Cache only the layer outputs or heads needed. Replace one corrupted-run activation with the corresponding clean-run activation.
  7. Compare recovery against no-op hooks, same-run patching, unrelated-position patching, and shuffled donor activations.
  8. Repeat on a held-out prompt family. Save raw per-example measurements, not just an averaged heatmap.

A useful normalized recovery is (patched_gap - corrupt_gap) / (clean_gap - corrupt_gap). Report raw gaps too. Exclude or separately analyze near-zero denominators; recovery can be negative or exceed one and is not a probability.

Current TransformerLens documentation recommends TransformerBridge and states that legacy HookedTransformer.from_pretrained was removed in version 4.0. It also distinguishes raw Hugging Face numerics from compatibility-mode transforms. Preserve the version expected by an ARENA notebook, or migrate deliberately and rerun parity checks.

Memory and compute planning

Weights are only one memory consumer. A single cached residual tensor needs approximately batch × sequence × width × bytes-per-element. At batch 1, length 256, width 768, float32, that is 0.75 MiB per layer. Twelve such tensors total 9 MiB, but caching multiple hook points, attention matrices, logits, gradients, SAE features and repeated donors increases this substantially. These are arithmetic estimates, not measured application peaks.

Use inference mode when gradients are unnecessary; attribution-gradient experiments must enable gradients. Start with batch 1, filter cached hook names, and move saved results to CPU. Training a large SAE is a later compute decision. NNsight supports remote execution, but check model access and resource terms before choosing it.

Research standards and common traps

  • A probe shows decodable information; it does not demonstrate causal use.
  • Zero ablation can create unnatural states. Compare mean, resample and matched counterfactual interventions where appropriate.
  • Activation patching depends on the corruption and metric; change both as robustness checks.
  • Direct logit attribution does not automatically account for later nonlinear computation.
  • SAE feature labels are hypotheses. Evaluate reconstruction error, sparsity, feature splitting, and the effect of interventions on the original model.
  • A circuit found on a narrow prompt distribution is not a general model explanation.
  • An attractive attribution graph is a starting point for tests, not proof.

First three actions

  • Read I01 and write one behavior-level research question.
  • Complete the first two relevant ARENA exercise sets in a pinned environment.
  • Produce a clean/corrupt activation-patching notebook with three negative controls.

Record completion with an evidence link and a helpfulness rating in a personal study log.