Automated TAE interpretation — status and proposed next work, 9 September 2026 Follow-up: [the consolidated things-to-do document](THINGS_TO_DO.md) supersedes this note's proposed work order. The subsequent four-report reconciliation found that FABLE already contains query-conditioned reader positives; recover and reconcile those experiments before treating that approach as new. The later multilingual pilot is a new execution, separate from the read-only audit described here. My recommendation is to build a small, falsifiable benchmark for automated TAE interpretation around the existing TRACE organisms and SONAR assets. The decisive question is whether an explanation helps predict unseen activations or intervention outcomes, and whether latent access improves those predictions over strong baselines under a stated budget. The existing program has substantial infrastructure and useful negative results, but has not established a robust automated-interpretation advantage on its corrected benchmark. This is an audit of local reports, selected result JSONs, and live file metadata on `4l`, plus a targeted primary-source literature search. I did not rerun training, load the retained checkpoints, reproduce all results, or establish publication readiness. Most experimental evidence dates from June–August; the latest tracked consolidation is August 14. September literature/site updates are newer than the experimental README. Treat recommendations below as proposals. **The four benchmark names refer to different objects.** | Artifact | What it evaluates | How I would use it | |---|---|---| | [SIEVE](/data/research/benchmarks/sieve_bench/README.md) | Sentence encoders, across reading, decomposition, construction, editing and generalization; 26 tasks | An encoder characterization suite; useful background for selecting substrates | | [GAUGE](/data/research/benchmarks/gauge/README.md) | Interpretability techniques on authored synthetic spaces with known answers | Fast ground-truth sanity tests and negative controls | | [TRACE / night8](/data/research/benchmarks/night8/NIGHT8_REPORT.md) | Techniques on trained miniature TAEs, with per-model certificates, baselines, confidence penalties and some SONAR transfer | Best starting point for the automated-interpretation benchmark | | [TAE-Bench, Campaign100 row 100](/data/research/tae-interp/experiments/campaign100/100-tae-bench-release/tae-bench/README.md) | Primarily the binding battery and reusable measurement checks, including cached SONAR embeddings | Reproduction/instrumentation package; its claims and controls need reconciliation with the later H1 report | The benchmark architecture is already largely present. A new name or another broad task list would contribute less than a clear experimental contract and a held-out evaluation. **What the automated-interpretation experiments established.** TRACE trained a pilot, canonicalizing organisms at different bottleneck widths, a count-aggregation organism, seed replicates, and a misreporting organism. The useful distinction is between factors observable in decoded output, factors readable from the latent but suppressed in output, factors merged together, and factors absent by construction. The count organism is particularly reusable: the output gives the total, while the latent can retain properties of the unordered count pair. Supervised readouts demonstrate some reachable headroom. The most informative result was a baseline correction. A method could inspect the original input while the reference method only inspected decoded output. That let input-derived labels look like a discovery about the latent. After stronger baselines and behavioral validation of merged claims, `distill_verify` had **premium 0.000, 95% CI [−0.167, +0.222]** on the v2 `basis_full+` comparison. I checked this directly in the [r9 result JSON](/data/research/benchmarks/night8/maker_r9/results/r9_v2.json). The interval does not establish exact equivalence; it establishes no statistically resolved positive advantage in that comparison. The [r9 report](/data/research/benchmarks/night8/maker_r9/MAKER_R9_REPORT.md) records the corresponding correction. Other recurring findings: surface-based methods confidently claim the wrong structure; abstention/calibration separates methods more clearly than discovery does; latent editing did not show a positive premium over strong string-edit baselines in the reported seed tests. These are results about the tested methods, tasks and budgets. The old report's “airtight fixed point” language exceeds what a finite adversarial test suite establishes. The early SONAR output-centric auto-interpreter still exists [locally under parascopes](/data/research/parascopes/layerwise/src/tae_autointerp_flipbook.py) and at `4l:/workspace/HOME/guest/parascopes_layerwise/src/tae_autointerp_flipbook.py`. The [running results](/data/research/tae-interp/docs/results/INTERP_RESULTS.md:261) report detection around 0.66 at perturbation dose α=20 and 0.71 at α=32. But high-dose interventions can collapse text into repeated letters and still receive excellent detection scores. These scores establish distinguishable effects under the old protocol, not faithful semantic explanations. Recovering a few hundred varied features for a stronger evaluation is more valuable than simply extending that atlas. **What the SONAR work established, with later corrections.** | Thread | Evidence worth retaining | Scope / correction | |---|---|---| | Role binding | H1 cross-construction plus vocabulary-holdout AUC: linear **0.509 [0.489, 0.529]**, MLP **0.495 [0.471, 0.519]** | Linear power check passes for an additive planted signal; MLP power check fails at the target effect. Genitive is also ~0.50 in this exact transfer setting. This is not proof that role information is absent under every readout. | | Japanese “exception” | Old ~0.696 result reproduces on its original lexicon, but becomes ~0.451 on a third lexicon | Withdraw the stable-generalization claim. The assay's arbitrary focal-label convention needs repair; failure of this assay cannot settle whether Japanese binding exists in another coding. | | Content-dependent order | H3: within a fixed word multiset, order AUC ~0.997; cross-pair global readout ~0.523 | Supports content-dependent readability. Real SONAR requires ~23 PCs for 90% of the measured order variance, rather than the toy model's ~4. These diagnostics do not uniquely identify a causal mechanism. | | Natural-text steering | H2: fluent-gated negation success ~0.70 at α=1; tense reaches ~0.85 at α=1.5–2 | Higher doses garble text. Above/below transfer is ~0.04 on natural, often nonspatial usages. Report applicability and content preservation alongside success. | | Round-trip monitor | H2 garble discrimination AUC ~0.845; a withholding tool was built | Threshold was calibrated on the same sweep. This is a useful prototype, not an independently validated detector of arbitrary semantic fabrication. | | Earlier binding localization | Valid retained token states and pooled depth profiles | The dedicated binding-death study ended **INSTRUMENT_FAILURE** after label alignment and post-LayerNorm subtraction problems. H3 reuse of those states does not retroactively validate the failed localization tests. | Sources: [H1](/data/research/tae-interp/experiments/consolidation/H1-binding-breaker/RESULT.md), [H2](/data/research/tae-interp/experiments/consolidation/H2-steering-buildup/RESULT.md), [H3](/data/research/tae-interp/experiments/consolidation/H3-theory-real-sonar/RESULT.md), [binding-death post-mortem](/data/research/tae-interp/experiments/binding_death/BINDING_DEATH_RESULTS.md). The [consolidation overview](/data/research/tae-interp/experiments/consolidation/CONSOLIDATION.md) is convenient but more confident than these detailed reports warrant. The SAE archive contains two distinct substrates: dictionaries over whole SONAR `z`, and dictionaries over the derived pooled component `c_pool`. Their results should not be silently combined. In the W40 `c_pool` sweep, two-seed mean invariant-feature counts stay around **1,200–1,400** across 8k–65k widths, while the measured invariant energy fraction falls **43.4% → 32.5% → 23.1% → 16.2%**. The result JSON itself flags possible undertraining. This is not an intrinsic information-capacity bound or a universal 16% semantic fraction. The earlier 273-atom/16% result also used a different paraphrase bank. See the [scale result](/data/research/tae-interp/experiments/w40/results/w40_scale_summary.json). The [vocabulary-invariance audit](/data/research/tae-interp/experiments/w40/VOCAB_INVARIANCE.md) found that among classifiable features, 51–55% passed its abstract rewording criterion and 18% were classified lexical. This is encouraging for controlled feature-description evaluation. Matched-topic controls were also run: their JSONs are in `tae-interp/experiments/w40/results/matched_topic_null_*`. The README still lists this as owed. Their stricter arms drop many pairs, so compare matched and ordinary nulls on the same retained subset before interpreting the smaller survivor count. **What is actually still on `4l`.** All paths below are relative to `4l:/workspace/HOME/guest/`. Sizes are approximate `du` measurements, not unique-data totals. [The live inventory](4l-inventory.json) records exact checkpoint paths, sizes and timestamps. | Directory | Approximate size | Reusable assets verified | |---|---:|---| | `night8/` | 8.0 GB | 10 `.pt` checkpoint files across the pilot, v1, v2, v3 and seed replicates; three saved r9 boards | | `campaign100/` | 16 GB | Experiment archive; local row directories supply reports, code and output records | | `consolidation/` | 94 MB | H1/H2/H3 code and outputs; H2 includes operators and judged natural-text edits | | `ladder/` | 13 GB | 28 final checkpoints, including a BART anchor; DAE/paraphrase/mix variants | | `binding_death/` | 7.8 GB | 52 cached layer `.npz` files across agent/patient and genitive; extraction/readout metadata | | `w40/` | 6.0 GB | Eight `c_pool` SAE checkpoints: 8k/16k/32k/65k × two seeds; two 131k whole-`z` SAE checkpoints | | `statspass/` | 560 MB | Binding/stimulus assets; `stimuli_v2/out/emb_cache_sonar.npz` exists (~28.8 MB) | | `real_parascope/` | 39 MB | Adjacent transfer experiment archive | | `parascopes_layerwise/` | 23 GB | Broader experiment code/assets, including the flipbook auto-interpreter | The old checkpoint manifest lists whole-`z` SAE checkpoints at several smaller widths, but only the two 131k files remain in its specified `w40/sae_curve_ckpt/` directory. Those paths must not be assumed runnable. The eight `c_pool` checkpoints are a separate, intact collection. File existence does not establish loadability; that remains a cheap next verification. At the live check, all four A4000s reported 0% utilization, with little allocated memory, and `/workspace` had about 490 GB free. This is a snapshot; the old note about a dead GPU and tight disk should not be treated as current machine status. No new experiments were launched. **The next experiments I would prioritize.** 1. **Freeze one benchmark contract and verify it can reward a real solution.** Start with TRACE, preserve its known-negative traps, and separate (a) predicting a latent feature's activation, (b) predicting an intervention's effect, and (c) practical uplift from latent access over text access. For (a), score descriptions on unseen positives, nearby negative concepts and rewordings. For (b), make the prediction before revealing the intervened decode. For (c), match text, sampling and compute access. Use privileged ground-truth methods as positive controls and corrupted explanations/random directions as negatives. A benchmark where every honest method must tie the baseline is not sufficient to measure progress. 2. **Run a compact method comparison on retained checkpoints.** A reasonable pilot is 100–200 stratified features: frequent/rare, lexical/rewording-robust, read-only/steerable, and selected negative controls. Compare a one-shot activation-example explainer; an output-centric explainer using perturbation pairs; a combined explainer; and an adaptive hypothesis-testing agent that proposes discriminating examples and revises its claim. Give the adaptive method an equal-budget nonadaptive comparator. Include random/PCA directions and the original text-only or label-mining baselines as appropriate to each track. These sample counts are design suggestions, not power calculations. 3. **Require measurable, held-out consequences.** Each explanation should specify an activation prediction, intervention prediction, domain of applicability, and confidence. Report prediction loss, false confident claims, coverage, collateral changes, and cost separately. Hold out whole propositions/templates and vocabularies, not just sentences; generate fresh private organism variants for final evaluation. Add seed uncertainty and a small blinded human audit of fluency/meaning labels. Before declaring an advantage, preregister a meaningful effect size, establish power for it, and verify it on a second trained substrate. 4. **Repair the binding question with a query-conditioned reader.** Reuse cached SONAR states. Ask `g(z, queried_entity)` whether that entity is the agent, instead of expecting an unconditional global direction to recover an arbitrary alphabetical designation on unseen nouns. Calibrate with both additive signals and planted role/filler rotations or tensor-product bindings, under matched distractors and vocabulary/construction holdout. A positive result would revise the global-axis interpretation; a negative result with adequate power would strengthen a clearly specified reader-family claim. Start with this cheap test before another architecture or scale sweep. 5. **Escalate to objective interventions only after the assay works.** The canonicalizing toy TAEs already show stronger relational readouts than SONAR under their own tests. Re-evaluate both with the same repaired assay, then compare reconstruction training against explicit relational supervision under matched capacity/data. The causal question is whether training makes relational information more accessible without sacrificing reconstruction; the existing cross-model contrast alone cannot isolate the cause. There is a subtle limit to “information beyond text”: for a fixed deterministic encoder, `z = E(text)` cannot add unrestricted Shannon information to the complete text. The meaningful targets are better predictions under constrained readers or budgets, and model-specific internal/causal facts. Matched models with similar outputs but different internal implementations would be a useful later test of the latter. Conditional probing (see reading list) provides a useful formal starting point. I would defer another broad 100-experiment campaign, larger SAE training solely to improve detection scores, and declaring the old negative a general impossibility theorem. The archive is large enough to answer the sharper questions first. The ledger's 0.993-style numbers are outputs of a heuristic evidence aggregation scheme over related experiments; they should not substitute for independent replication or be presented as calibrated probabilities of scientific truth. **A focused reading order.** These links were checked against primary pages during this audit; this is a targeted reading list rather than a comprehensive literature review. | Priority | Paper | What to take from it | |---|---|---| | 1 | [CHIVE: Would This Change Your Answer?](https://alignment.anthropic.com/2026/chive/) — August 2026 | Evaluate explanations by predictions about measured counterfactuals. Its tested activation-reading tools did not outperform a transcript-only predictor; directly relevant to TRACE's baseline correction. | | 2 | [Pitfalls in Evaluating Interpretability Agents](https://arxiv.org/abs/2603.20101) — March 2026 | Agents can reproduce published explanations through guessing or memorization. Read before designing the private test split and explanation scoring. | | 3 | [Automatically Interpreting Millions of Features in Large Language Models](https://arxiv.org/abs/2410.13928) — 2024 / revised 2025 | Practical automated explanation pipeline and several scoring methods, including intervention scoring. Baseline implementation reference. | | 4 | [Enhancing Automated Interpretability with Output-Centric Feature Descriptions](https://arxiv.org/abs/2501.08319) — ACL 2025 | Activating inputs and causal output effects support different descriptions; combining them improves the tested evaluations. Closest fit to TAE encode/decode access. | | 5 | [Automated Interpretability and Feature Discovery in Language Models with Agents](https://arxiv.org/abs/2605.01555) — May 2026 | A concrete competing-hypotheses and targeted-controls approach; useful for the adaptive method arm. Reported improvements are on its tested LLM features, not established for TAEs. | | 6 | [Are Sparse Autoencoder Benchmarks Reliable?](https://arxiv.org/abs/2605.18229) — May 2026 | Read the reseed-noise, ground-truth correlation and discriminability audit before inheriting SAE metrics. It finds serious problems with TPP/SCR at their canonical settings. | | 7 | [Evaluating Neuron Explanations: A Unified Framework with Sanity Checks](https://arxiv.org/abs/2506.05774) — ICML 2025 | Your metric must worsen when concept labels/explanations are corrupted. This should become a benchmark regression check. | | 8 | [SynthSAEBench](https://arxiv.org/abs/2602.14687) — revised July 2026; [InterpBench](https://arxiv.org/abs/2407.14494) — NeurIPS 2024 | Complementary ground-truth substrates: realistic synthetic feature distributions versus trained semi-synthetic transformers with known circuits. Use them to position GAUGE/TRACE. | For the binding/theory branch, add [Conditional probing: measuring usable information beyond a baseline](https://aclanthology.org/2021.emnlp-main.122/) and [Discovering the Compositional Structure of Vector Representations with Role Learning Networks](https://arxiv.org/abs/1910.09113). The former sharpens the readout-relative claim; the latter is directly relevant prior work on discovering role/filler structure in sentence representations. For SAE null design, [Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?](https://arxiv.org/abs/2602.14111) provides concrete random baselines that rival trained SAEs in its experiments. If reading time is limited, start with CHIVE, Pitfalls, and the output-centric paper. Together they address the research question, the evaluation failure modes, and the most natural TAE-specific method.