ai gen
TAE research / completed analysis

SAE methods and implementation audit

← Back to the plan and kanban

SAE state of the art — 9 September 2026

Literature research delegated to the sae_sota subagent, combined with a local implementation audit. No training or evaluation reruns were performed.

The practical recommendation for this research program is Matryoshka BatchTopK as the next challenger, with ordinary BatchTopK retained as the baseline. No universal winner is established, particularly on sentence embeddings. BatchTopK controls sparsity; Matryoshka supplies nested reconstruction objectives. They combine naturally.

Approach Role Qualification
BatchTopK Strong simple reconstruction/sparsity baseline; variable active-feature count per example Calibrate a fixed inference threshold; training batch selection is not the inference rule
Matryoshka + BatchTopK Encourage broad concepts and refinements at different dictionary prefixes Feature organization is the motivation; not a universal reconstruction advantage
Matryoshka + JumpReLU Alternative nested-loss recipe with learned activation thresholds Used by Gemma Scope 2, alongside additional sparsity/frequency penalties
Ensembles Reuse independent dictionaries to improve coverage and stability Match total dictionary width, active-feature budget and cost; superiority to Matryoshka is not established
Group Bias Adaptation Optional challenger using target activation-frequency groups Synthetic feature-recovery guarantees do not establish recovery on SONAR

Primary sources: BatchTopK, Matryoshka, Gemma Scope 2 report, Ensembling SAEs, GBA conference paper.

Matryoshka's paper reports improvements in absorption and several downstream measures while automated interpretation remains comparable to BatchTopK. Nested objectives encourage early general features and later specialization. Its implementation reported roughly 50% extra training time for five nested dictionaries. Transfer of these benefits to pooled SONAR vectors needs measurement. Paper

The May 2026 benchmark reliability audit is essential context: canonical TPP and SCR settings fail reliability tests, and other metrics also have substantial noise. The audited sae-probes sparse-probing variant performed comparatively well, but closely related architectures remain difficult to distinguish. Older multi-metric leaderboard wins therefore do not establish a definitive current ranking.

The local program has already tried Matryoshka. Its saved h16384/k128 model reports FVU 0.3104, with prefix FVUs 0.4052 / 0.3447 / 0.3104 for 1024 / 4096 / 16384 features. The running report puts plain BatchTopK at approximately 0.311. These are encouraging exploratory reconstruction ties, not a controlled feature-quality verdict. Saved metrics

Two code details need attention before reuse:

  • TopKSAE.encode always performs batch selection when BatchTopK is enabled, including evaluation. Both published BatchTopK and Matryoshka recipes replace that with a calibrated global threshold at inference. Otherwise a sentence's feature activations depend on its batch companions. BatchTopK recipe
  • The local Matryoshka training loop averages prefix reconstruction losses but does not call AuxK. Global BatchTopK before slicing prefixes is consistent with the published formulation; omitted AuxK and inference calibration are the main identified differences. Averaging instead of summing also matters when balancing auxiliary objectives.

The recommended next comparison is deliberately small:

  1. Calibrate the retained BatchTopK models on a separate split; report held-out actual mean L0 and its distribution, reconstruction, dead features and batch invariance.
  2. Train a faithful Matryoshka BatchTopK comparator at one existing width and sparsity, with the same examples/normalization and at least two seeds. Preserve the plain baseline.
  3. Measure held-out activation prediction, paraphrase consistency, hierarchy/absorption controls, intervention effects, decoded fluency and collateral changes. Include confidence intervals and corruption/random controls. Expand the sparsity sweep if the first comparison is informative.
  4. Use retained seed pairs for an ensemble pilot before another width increase; match both total width and total active features.
  5. When adding another text autoencoder, cross encoder choice with SAE choice as separate experimental factors. Re-encode examples and train new dictionaries; equal dimensionality does not establish coordinate compatibility. The subsequent access audit and user decision park OmniSONAR; see the consolidated plan for accessible alternatives.

These are proposed experiments, not literature-established outcomes. An old local JumpReLU loss does not rule out a better tuned current JumpReLU recipe.

For later investigation, Matching-Pursuit SAE offers an iterative residual encoder and a useful challenge to shallow feature recovery on correlated data. Its computation and stability need explicit comparison. Rational SAE, a June 2026 workshop paper, is a fine-tuning lead rather than an established replacement. Ensembling's paper explicitly leaves extensions to Matryoshka for future work; its reported gains should not be generalized to that comparison.

Suggested reading order: Matryoshka, BatchTopK, the benchmark reliability audit, Gemma Scope 2, Ensembling SAEs, then Matching-Pursuit SAE. The first three are sufficient to design the immediate experiment.