When is an Interpretability Feature Faithful?
收藏资源简介:
The repository contains the code and data relevant to the article. All details are provided in the article. Abstract: Mechanistic interpretability attributes a network's behaviour to internal units by intervening on them, but their scores are not inferential. An ablation ratio or a patching recovery is a point estimate with no controlled error, sensitive to the intervention that produced it, and read one unit at a time from a dictionary of thousands of correlated units. We instead make a pre-specified operational claim the object of a test: a unit, or a group of units, is certified when its ablation and installation contrasts, each passed through a bounded encoding and averaged over rounds, are both positive in the claimed direction, relative to named interventions and an input population. Each directional hypothesis is tested with a betting test supermartingale whose wealth is an e-value valid at any stopping time; their minimum tests the conjunction; and e-BH carries the locally stopped e-values to a false-discovery-rate guarantee over the declared candidate family under arbitrary dependence. Group certificates are exactly invariant to changes of basis within the tested span; per-coordinate certificates are not. Simulations with exact nulls are consistent with the type-I and false-discovery guarantees. On GPT-2, our proposed method, FACE, certifies four of the forty-eight screened features in its primary indirect-object-identification family, the same four when the pool is widened to one thousand, and the S-inhibition head set while declining the primary name-mover heads; on Gemma-2-2B, separately screened families abstain at layers 12 and 22 and certify one layer-18 feature. A certificate is relative to its intervention, population, and candidate family, and non-certification is not evidence of causal irrelevance.



