[Proposal] SVD Circuits: singular-vector decomposition of a head's QK/ OV into causally-validated subfunctions
Los mantenedores suelen responder en 1 día
@janmenjayap ya está trabajando en esto.
Desde el 10/9/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Proposal
Component-level circuit analysis (direct_logit_attribution, path patching, head detectors) treats an attention head as an indivisible unit. But a single canonical head — e.g. a GPT-2-small IOI name-mover — empirically superposes multiple subfunctions on distinct low-rank directions of its QK and OV maps. This proposes svd_circuits, a self-contained tool in transformer_lens/tools/analysis/ that takes the SVD of a head's QK (W_Q W_Kᵀ) and OV (W_V W_O) matrices, exposes each orthogonal singular direction as an interpretable sub-component, and — crucially — causally gates every claimed subfunction by patching activations along that individual singular direction. This is a finer-than-component decomposition that resolves within-head superposition component-level tools miss.
Capabilities:
- Per-head QK/OV SVD via
FactoredMatrix— get singular values/vectors ofW_Q W_KᵀandW_V W_Owithout materializing thed_model × d_modelmatrix. - Direction → vocab readout — project each OV singular direction into the unembedding basis (reusing
SVDInterpreter) and each singular direction into logit space viadirect_logit_attribution, so a direction gets a human-readable "what tokens does this subfunction move" signature. - Activation projection / attribution — project cached per-head activations onto the singular directions to measure how much each subfunction fires on a given prompt (per token/position).
- Causal validation (patch-along-a-direction) — reconstruct the head output using only a chosen singular subspace (or ablate a single direction) via
generic_activation_patch, and report the behavior change (e.g. IOI logit-diff), so a subfunction claim is accepted only if its direction is causally load-bearing. - Degeneracy guard — detect near-equal singular values (rotation/degeneracy ambiguity) and refuse to attribute individual directions inside a degenerate block, reporting the block as a subspace instead.
Status: PR1 (#1768) and PR2 (#1775) merged; PR3 in progress
Suggested labels: enhancement, tooling, TransformerBridge, complexity: moderate
Motivation
The existing TL baseline for "what does this head do" is direct_logit_attribution (transformer_lens/tools/analysis/direct_logit_attribution.py) and head_detector — both operate at head granularity. The IOI work established name-mover / S-inhibition / duplicate-token heads as the atomic units of the circuit. The claim in Beyond Components is that this granularity is too coarse: an IOI name-mover head is not one function but several, packed onto near-orthogonal low-rank directions of its QK/OV weights, and only an SVD-based decomposition surfaces them. This tool lets a TL user go inside a canonical head and attribute/patch its subfunctions — the natural next rung below component-level circuit analysis.
Why no maintained TL implementation exists. TL already ships the two hardest primitives — FactoredMatrix (low-rank OV/QK SVD/eigendecomposition without materializing the full matrix) and SVDInterpreter (projects OV/w_in/w_out singular vectors into vocab) — but neither closes the loop from "singular direction" to "causally-validated subfunction." SVDInterpreter is a static weight-inspection utility (its own docstring warns SVD directions are not reliably interpretable); it never touches activations or patching. The missing piece is the activation-projection + causal-patch layer, which is exactly what makes the decomposition trustworthy.
Implementation status elsewhere. The authors' research repo exists — github.com/Exploration-Lab/Beyond-Components (~7 commits; IOI / GP / GT examples) — but it is not maintained or TL-integrated. This would be the first TL-native implementation. This is not SAE/dictionary-learning territory (no learned features, no overcomplete basis) — it is a closed-form weight decomposition, so it sits inside TL rather than at the SAELens boundary.
Pitch
Add svd_circuits as a self-contained analysis tool under transformer_lens/tools/analysis/. It builds on FactoredMatrix for the weight-space SVD, SVDInterpreter for vocab readout, ActivationCache/direct_logit_attribution for activation projection and logit signatures, and generic_activation_patch for the mandatory causal gate.
Proposed API
Names adjustable to maintainer preference.
from transformer_lens.model_bridge import TransformerBridge
from transformer_lens.tools.analysis import svd_circuits
# TransformerBridge is the supported model system
model = TransformerBridge.boot_transformers("gpt2-small")
model.enable_compatibility_mode() # folds final LN into W_U — required for the OV vocab/logit readout
# 1. Decompose a head's QK and OV into singular directions (FactoredMatrix under the hood).
decomp = svd_circuits.decompose_head(model, layer=9, head=9, which=("QK", "OV"))
decomp.OV.S # singular values, sorted desc
decomp.OV.rank_report # (idx, sigma, sigma/sigma_max, is_degenerate) per direction
# 2. Readout: each OV singular direction -> top tokens in the unembedding basis.
decomp.OV.vocab_readout(k=10) # wraps SVDInterpreter.get_singular_vectors
decomp.OV.logit_signature(model, prompt=ioi_prompt) # via direct_logit_attribution
# 3. Attribution: how much does each direction fire on this prompt?
proj = svd_circuits.project_activations(model, decomp, prompt=ioi_prompt) # [pos, direction]
# 4. Causal gate: keep only directions {0,3}, measure IOI logit-diff change.
result = svd_circuits.patch_along_directions(
model, decomp.OV, keep=[0, 3], prompt=ioi_prompt,
metric=ioi_logit_diff, # any callable(logits) -> scalar
)
result.delta_metric # behavior change attributable to that subspace
result.gated # True only if |delta| exceeds the causal threshold
Design (algorithm)
For a head (layer ℓ, head h):
- Build factored maps.
OV = FactoredMatrix(W_V[h], W_O[h])(shaped_model × d_model, rank ≤d_head);QK = FactoredMatrix(W_Q[h], W_K[h].T). Both stay factored — never materialized_model². - SVD.
U, S, V = M.svd()(M == U @ S.diag() @ V.transpose(-2, -1);UandVare both[d_model, rank], with the i-th singular direction in columni—U[:, i],V[:, i]). Right singular vectorsV[:, i]are the input directions (which residual-stream directions the subfunction reads); left singular vectorsU[:, i]are the output directions. Access them via the.Vproperty —.Vhis a deprecated alias that returns the same tensor as.V(it emits aDeprecationWarning), not its Hermitian transpose; do not introduce a new call site that relies on it. - Degeneracy guard. Flag any run of singular values with
σᵢ/σᵢ₊₁ − 1 < ε(defaultε = 1e-2) as a degenerate block: individual directions inside it are rotation-ambiguous, so attribute the block as a subspace, not per-direction. - Vocab readout (OV). For each output direction
U[:, i], project through the (LN-folded) unembedding — reuseSVDInterpreter.get_singular_vectors("OV", ℓ, head_index=h)— to get its top-token signature. - Activation projection. Cache the head's per-position input (residual stream into the head; for OV attribution use
stack_head_results/ the head's value stream). Project ontoV[:, i]to get a[pos]firing coefficient per directioni. - Logit signature. Feed the rank-1 reconstruction
σᵢ · U[:, i] V[:, i]ᵀof the head output throughdirect_logit_attributionto get each direction's signed logit effect on the task metric. - Causal patch-along-direction. Using
generic_activation_patch, replace the head's output (atblocks.ℓ.attn.hook_z/hook_result) with its projection onto the chosen singular subspacespan(U[:, keep])(or zero out a single direction for ablation), rerun, and recordΔmetric. A subfunction is accepted only if its direction/subspace produces a causalΔmetricabove threshold — the decomposition alone never suffices.
Reuse map:
| Sub-step | TL primitive | file:symbol |
|---|---|---|
| Factored QK/OV, low-rank SVD/eigen | FactoredMatrix (.svd(), .U/.S/.V, .eigenvalues) |
transformer_lens/FactoredMatrix.py:FactoredMatrix |
| OV direction → vocab readout | SVDInterpreter |
transformer_lens/SVDInterpreter.py:SVDInterpreter.get_singular_vectors |
| Per-head activations for projection | ActivationCache |
transformer_lens/ActivationCache.py:stack_head_results (:957) |
| Direction → logit effect | direct_logit_attribution |
transformer_lens/tools/analysis/direct_logit_attribution.py:direct_logit_attribution |
| Bridge LN-folding guard (reuse pattern) | _validate_bridge_compatibility |
transformer_lens/tools/analysis/direct_logit_attribution.py:_validate_bridge_compatibility |
| Characterize resulting subfunction | head detector | transformer_lens/head_detector.py |
| Causal patch-along-direction | generic_activation_patch + setter |
transformer_lens/patching.py:generic_activation_patch |
| Weights / unembed / final LN | W_Q/W_K/W_V/W_O, model.W_U |
Bridge tl_parameters() |
Main correctness risk (named): interpreting individual singular directions under rotation/degeneracy ambiguity
SVD directions are unique only when singular values are distinct; equal (or near-equal) singular values leave the corresponding singular subspace defined only up to an arbitrary rotation, so any "direction 3 = the surname subfunction" claim inside a degenerate block is meaningless. Compounding this, "the SVD direction is interpretable" is not guaranteed even for well-separated directions (SVDInterpreter's own docstring flags numerical instability and cross-device inconsistency). Mitigation, not a warning: (a) the degeneracy guard (step 3) raises/refuses per-direction attribution inside a near-degenerate block and downgrades it to a subspace; (b) every claimed subfunction must pass the causal patch-along-direction gate (step 7) — a direction that is not causally load-bearing is never reported as a subfunction, regardless of how clean its vocab readout looks. The tool's contract is "causal evidence gates interpretive claims," not "SVD is interpretable."
Validation plan (falsifiable)
- Analytic SVD oracle. On a tiny synthetic head with hand-set
W_V/W_O(andW_Q/W_K), the tool'sS,U,Vmatchtorch.linalg.svdof the explicitly materialized matrix toatol=1e-5(sign/degeneracy handled). - Reconstruction fidelity. Summing all rank-1 direction reconstructions reproduces the full head output (via the cache) to
atol=1e-4; keeping top-kdirections drives abs(Δmetric) toward zero (non-increasing)slow - Degeneracy guard fires. A synthetic head with two exactly-equal singular values → the guard flags the block and per-direction attribution raises; subspace attribution still works.
- Causal gate discriminates. Ablating a high-
σ, causally-relevant direction changes the IOI logit-diff materially; ablating a null-space / random direction does not (baseline reported alongside). - Paper sanity check (slow). On
gpt2-smallhead L9H9 (name-mover), the tool surfaces ≥2 causally-gated OV subfunctions whose vocab readouts qualitatively match the paper's name-mover split; reported as a qualitative sanity check, not a pinned numeric threshold.
Scope of the vertical slice
Per-head QK/OV SVD + direction→vocab readout + direction→logit signature + one causal patch-along-direction validation, on gpt2-small IOI. Defer full multi-head circuit assembly and automated subfunction labeling. Ships as 3 sequential PRs, not one — base off dev:
-
feat(svd_circuits): per-head QK/OV SVD with degeneracy guard—decompose_head,HeadSVD, degeneracy guard. Pure weight-space, no public export yet. -
feat(svd_circuits): singular-direction readout, projection, and causal patch gate— vocab/logit readout,project_activations,patch_along_directions, public exports, Bridge compatibility-mode validation,gpt2-smallIOI integration test. -
docs(svd_circuits): add demo notebook, slow name-mover parity, and tool docs— Reflect the combined demo+docs commit,SVD_Circuits_Demo.ipynb, docs section.
Files:
- New:
transformer_lens/tools/analysis/svd_circuits.py - Export:
transformer_lens/tools/analysis/__init__.py - Unit:
tests/unit/tools/test_svd_circuits.py(synthetic head; checks 1–4, no HF download) - Integration:
tests/integration/test_svd_circuits.py(gpt2-small, IOI patch-along-direction end-to-end) - Oracle-parity (slow):
tests/integration/test_svd_circuits_oracle_parity.py(check 5,@pytest.mark.slow) - Demo:
demos/SVD_Circuits_Demo.ipynb(nbval)
Follow-up work
Deferred to a separate tiered issue, filed after PR3 merges: multi-head circuit assembly, automated subfunction labeling, QK-side pattern attribution, more models.
Model coverage
- CI reference / oracle: tiny synthetic head (no download) for unit correctness;
gpt2-smallfor the integration + slow paper sanity check (CI-cacheable, canonical IOI). - Demo model:
gpt2-small. - Honesty caveat: the paper's headline subfunction taxonomy is scale- and model-dependent;
gpt2-smallreproduces the mechanism (a name-mover splits into causally-distinct directions) but not necessarily the exact subfunction inventory of larger models. Baselines (random-direction ablation) are reported next to every headlineΔmetric.
Alternatives
- Stay at head granularity (
direct_logit_attribution/ head detectors). Rejected: cannot resolve within-head superposition — the whole point. - Use
SVDInterpreteralone. Rejected: static weight readout with no activation projection and no causal gate; its own docstring disclaims interpretability of the directions. - Sparse autoencoders on head outputs (SAELens). Rejected here: needs training, an overcomplete learned basis, and lives outside TL;
svd_circuitsis closed-form, training-free, and reuses TL primitives. - Full path-patching over a rank-1 direction grid. Rejected as the first slice: combinatorially expensive and premature before the single-direction gate is validated; folded into Tier 2 of the follow-up issue.
Correctness oracle
Brute-force / analytic reference (no external frozen impl exists). Primary oracle: torch.linalg.svd on an explicitly materialized synthetic head, compared to the FactoredMatrix-based path within threshold (atol=1e-5), plus a reconstruction-fidelity check. Secondary (slow, @pytest.mark.slow): the gpt2-small name-mover paper sanity check, treated as qualitative because the paper ships no numeric reference to pin against.
Additional context
- Paper: Areeb Ahmad, Abhinav Joshi, Ashutosh Modi (IIT Kanpur), "Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits," arXiv 2511.20273, NeurIPS 2025 poster (poster 119702, OpenReview
7UbXEQNny7). - External implementation status: authors' research repo
Exploration-Lab/Beyond-Componentsexists (research-grade, ~7 commits, 4 stars, notransformer_lensdependency); this would be the first TL-native implementation. Paper nuance: it augments weight matrices with biases folded in and applies SVD uniformly to attention and MLP — this slice covers attention QK/OV (OV logit-facing) only. - Artifact sources: none required — the decomposition is computed from model weights at call time; no downloaded lens/SAE artifacts, so no registry entry is needed for the vertical slice.
- Algorithm note for reviewers: OV singular directions are attributed through the unembedding (output side); QK singular directions describe which residual directions the head compares (attention-pattern side) and are attributed via the pattern, not the unembedding — the vertical slice focuses on OV for logit-facing claims and includes QK SVD as a readout only. The causal gate is what separates this from a pretty-picture SVD tool; keep it mandatory.
- Substrate history: SVD convention fix #341 (closed) → #1300 (merged) (
torch.linalg.svd;.Vhdeprecated in favour of.V); #328 (closed, 2023, "SVD tests fail on GPU") is direct evidence for the cross-device instability the degeneracy guard + causal gate backstop. - Provisional-citation flags: arXiv id
2511.20273and NeurIPS-2025 venue are confirmed (no open citation flags).
Checklist
- I have checked that there is no similar issue in the repo (required).
- Checked no similar issue / tool exists (
SVDInterpreteris static weight readout only; no causal singular-direction tool intools/analysis/). - Scope limited to a vertical slice (single head, OV logit-facing, one causal gate); follow-up tiered separately.
- Correctness oracle identified (analytic
torch.linalg.svd+ reconstruction; slow qualitative paper check). - Targets
TransformerBridge(the supported TL 3.x path) via sharedActivationCache - Claim-in-comments before starting a tier item.
- Lenguaje dominante
- Python
- Estrellas
- 3.9k
- Forks
- 708
- Merge medio
- 1 d 17 h
- PR fusionados (30 d)
- 70
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de TransformerLensOrg/TransformerLens
-
[Proposal] RoBERTa masked-LM adapter for TransformerBridgePosiblemente ocupada @Canonik la tomó hoy. Abiertocomplexity-moderate new-architecture TransformerBridge
TransformerLensOrg/TransformerLens#1870 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Proposal] Backward Lens: support gated MLP gate/ up/ down gradient factorsPosiblemente ocupada @janmenjayap la tomó hace 8 días. Abiertocomplexity-moderate enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1832 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Proposal] Sparse probing: optional groups argument so rows from one prompt can't straddle the splitPosiblemente ocupada @lorenzozanee la tomó hace 13 días. Abiertocomplexity-simple enhancement help wanted TransformerBridge
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
TransformerLensOrg/TransformerLens#1813 ·
Los mantenedores suelen responder en 1 día
-
[Bug Report] _BLOCK_LIST_ATTRS hardcoded name list silently drops Raven's blocks from composition-score / head-label analysisPosiblemente ocupada @LightWork666 la tomó hace 17 días. Abiertobug complexity-moderate TransformerBridge
TransformerLensOrg/TransformerLens#1791 · 2 comentarios · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Proposal] Relevance Lens (R-lens): a RelP/ LRP-based transport-matrix estimator for Jacobian Lens (J-lens) fit, readout, and interventionPosiblemente ocupada @janmenjayap la tomó hace 31 días. Abiertocomplexity-high enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1755 · 6 comentarios · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de TransformerLensOrg/TransformerLens
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
[BUG] Multi-day events show "Ended" while still in progressPosiblemente ocupada @tarunagnihotri534 la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
data-umbrella/du-event-board#231 · 2 comentarios ·
-
avl_automation: the generated control surface block isn't valid XML (typo in avl_out_parse.py)Posiblemente ocupada @brksol la tomó hoy. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
PX4/PX4-gazebo-models#164 ·