Ranking a model’s parts wins a benchmark for tracing its answers

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

TL;DR On a hidden test of finding which parts drive a model’s answers, a new ranking method averaged 5.6 (higher is better) against 1.95 for the runner-up. A variant traced refusals to 1% of a chat-tuned Llama model’s weights, a tool for auditing what tuning changed.

The problem

When a chatbot answers, which of its internal parts did the work? Figuring that out is called attribution, and each existing tool fails somewhere. You can swap out one part and watch the answer, but that costs a separate run per part. Calculus shortcuts score every part at once, yet may point at the wrong ones. Training an on-off switch per part works but is finicky, and must be redone for every explanation size.

The idea

The team’s method, Matryoshka Attribution, learns one score per part instead of one switch per size. Each training step draws a random budget and keeps only that many top-scoring parts, like a coach fielding a randomly sized team at each practice. Benched parts get their values from a slightly altered prompt, and scores are nudged so the model still gives its original answer. Over many practices the coach settles on one ranking, where each smaller lineup nests inside a bigger one like Russian dolls. Because every size gets practice, the authors explain, one training run yields a ranking that works at all of them.

What they showed

  1. Top of the leaderboard. On the Mechanistic Interpretability Benchmark, which checks whether a method’s top-ranked parts alone reproduce a model’s answers, it averaged 5.6 on the hidden test against 1.95 for the runner-up.
  2. Refusals sit in a sliver of weights. A second version ranks weights, the numbers a model learns in training. Resetting the top 1% of Llama 3.1 8B Instruct’s weights to pre-tuning values raised an AI judge’s harmful-compliance score from 2.6 to 84.0; math and knowledge scores barely moved.
  3. Cheaper than sweeping. Rival switch methods trained at many sizes cost about 50 times as much and still scored below it.

Why it matters

For researchers probing models, one run now ranks parts at every size. The authors pitch it as a step toward making model explanation a problem you solve by training. For people who audit open-weight models, the authors say the refusal result shows the risk of releasing an unguarded pre-tuning model beside its safety-tuned version. Anyone holding both might undo refusals by resetting a sliver of weights.

The catch

The benchmark’s main score has a blind spot the authors name themselves: it underrates performance with small sets of parts. On their stricter score, the method was only close behind a calculus shortcut on some kinds of parts, so the leaderboard gap overstates its edge there. The refusal work took its reward from the AI judge of StrongREJECT, a set of harmful requests, and was graded on other prompts from it, though a separate refusal test agreed.

Who wrote it

Aryaman Arora and six Stanford University colleagues, among them Christopher Potts. Arora and Potts co-authored “Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability” (2023), and the new method is grounded in causal abstraction.