ai gen
Research plan / 9 September 2026

What to do
next.

Turn the TAE archive into a benchmark that can reward useful explanations. Reuse what is trained, test the multilingual idea, and compare methods only after the controls work.

12 open tasks2 completed artifacts4 source reports

The multilingual pilot is done. Direct French→Chinese decoding works on simple role and number contrasts. The small study does not establish an internal English bottleneck or a better translation route. Read the results and decoding-sensitivity check →

These are the new research tasks. The Campaign100 board preserves the earlier campaign. Select a card to open its complete protocol; statuses are a recorded snapshot, not a personal checklist.

14 cards · 12 open tasks and 2 completed artifacts

Ready next4

Can start with the current evidence and assets.

T01 / P0Recover a reproducible set of assets

Can another run recover the exact representations and outputs used by the first comparison?

No prerequisite task

T02 / P0Reconcile the scientific claims

Which claims still hold once the earlier FABLE positives, later H1–H3 checks and corrected benchmark baselines are considered together?

No prerequisite task

T05 / P1Extend the multilingual pivot experiment

What information changes when a French sentence vector is decoded directly into Chinese or passed through English text or an English-reference vector map?

No prerequisite task

T09 / P2Evaluate an accessible second autoencoder

Do the findings and interpretation methods survive on an independently built, usable sentence autoencoder?

No prerequisite task

Needs prerequisites8

Work begins after the linked task gates pass.

T03 / P1Freeze and validate the benchmark

Can the evaluation distinguish a useful interpretation from a plausible story or a surface shortcut?

Needs T01

T04 / P1Compare automated explanation methods

Does an explanation predict new activations and intervention effects better than matched baselines?

Needs T03, T06

T06 / P1Calibrate SAE inference and compare recipes

Does a faithful Matryoshka recipe yield more useful features than plain BatchTopK at comparable cost and sparsity?

Needs T01

T07 / P2Repair binding measurement before training

Which reader can recover which role information under real vocabulary and construction transfer?

Needs T02

T08 / P2Separate decoder failures from monitor failures

Can a monitor abstain on semantic errors at useful coverage, and how much does decoder training change the answer?

Needs T01, T02

T10 / P2Run only claim-changing follow-up controls

Which remaining uncertainty could change a central paper claim with the least new work?

Needs T02

T11 / P2Finish a paper with defensible scope

What is the strongest coherent contribution supported by the reconciled evidence?

Needs T02

T12 / P2Package a reproducible release candidate

Can someone reproduce the benchmark contract and its controls without the original workspace?

Needs T03, T11

In progress0

Only experiments or analysis actually underway.

No tasks in this lane.

Completed2

Finished artifacts; their limitations still apply.

D01 / DoneConsolidate the four project reports

Combined the roundup, paper, literature review and campaign board with the status audit, SAE review and 4l inventory. Scientific claims still require T02 reconciliation.

Completed in this session

D02 / DoneRun the multilingual pilot

24 French/English/Chinese triples; nine direct routes and three French pivot routes, with greedy decoding and beam size 5. No fitted vector map or independent bilingual scoring yet.

Completed in this session

Expanded task protocols

Each task states the question, procedure, controls, output and decision rule. Effort bands describe the work, not a reserved GPU schedule.

Task data · JSON
T01 · P0Recover a reproducible set of assets

Can another run recover the exact representations and outputs used by the first comparison?

The archive contains useful trained models, but a file inventory alone does not establish loadability. Whole-z and c_pool dictionaries answer different questions, and some smaller whole-z checkpoint paths are missing.

Prerequisites: None — can start now
Compute: CPU audit + short checkpoint smoke

Procedure

  1. Select one TRACE organism, one c_pool dictionary, one whole-z dictionary and the retained SONAR encoder/decoder. Do not attempt to validate every archived checkpoint at once.
  2. Record model names, checkpoint paths and hashes, dimensions, normalization, tokenizer, package versions, seeds, dataset provenance and split identifiers. Resolve paths inside the chosen environment rather than copying historical absolute paths.
  3. Load each selected artifact and reproduce a saved reference prediction or reconstruction. Check tensor shapes, finite activations and the expected representation type.
  4. Write a portable manifest with an explicit unavailable/excluded entry for anything that cannot be recovered. Link each reproduced output to its source result.

Required controls

A successful deserialization is insufficient. Compare a known reference output and a deliberately mismatched configuration; the latter must fail clearly.

Measurements

Reference-output agreement, numerical tolerance, checkpoint coverage, missing assets and peak memory for the smoke run.

Deliverable

asset_manifest.json, pinned environment notes, and smoke_results.json with a runnable command.

Decision / stopping rule

Start the benchmark comparison only with artifacts that reproduce their reference behavior. Missing historical checkpoints are exclusions, not a reason to silently substitute a new model.

First action: Read 4l-inventory.json and choose the four representative assets. Estimate the cost from their actual file sizes before loading.
T02 · P0Reconcile the scientific claims

Which claims still hold once the earlier FABLE positives, later H1–H3 checks and corrected benchmark baselines are considered together?

The four source pages describe different snapshots. Their broadest headlines can disagree even when their underlying task definitions differ.

Prerequisites: None — can start now
Compute: Reading + cached result analysis

Procedure

  1. Create one row per claim with its exact estimand, substrate, reader family, available query information, vocabulary/construction split, power gate and independent supporting result.
  2. Map the existing query-conditioned reader and decoder-parser positives against H1. Determine whether each difference is a changed task, changed access, failed instrument or a genuine contradiction.
  3. Reconcile W40 matched-topic results using the same valid subset and denominator. Mark undertrained reconstruction results and keep c_pool separate from whole-z.
  4. Update the proposed abstract and board wording: separate gross from subtle fabrication, post-hoc capacity estimates from frozen tests, and observational geometry from causal localization.
  5. Only then rerun the historical evidence ledger with the corrected support/refutation inputs. Preserve the old version and label the aggregation heuristic.

Required controls

Inspect primary result JSONs and code for each disputed number. Do not resolve conflicts by selecting the newest report or the highest confidence label.

Measurements

Claims with traceable evidence, unresolved contradictions, instrument failures, corrections and remaining replication requirements.

Deliverable

claims_reconciled.csv plus a concise change log and proposed paper/board wording.

Decision / stopping rule

No claim is formally promoted simply because bookkeeping is complete. Every promoted claim needs its stated evidence and power requirements.

First action: Start with binding, the 460-bit estimate, the 16% atom-energy claim and the corrected TRACE premium.
T03 · P1Freeze and validate the benchmark

Can the evaluation distinguish a useful interpretation from a plausible story or a surface shortcut?

TRACE previously rewarded unequal input access. A benchmark where the best permitted method must tie a strong baseline cannot measure the progress we want.

Prerequisites: T01
Compute: CPU controls + small inference pilot

Procedure

  1. Specify three tracks separately: unseen activation prediction, unseen intervention-outcome prediction, and practical uplift from latent access under a matched budget.
  2. Choose a small TRACE organism set. Use GAUGE to validate the measurement instrument and SIEVE/TAE-Bench components for stimulus generation and encoder capability checks.
  3. Freeze information access, training/validation/test boundaries, the unit of independence, the primary score, a minimum useful effect, and the stopping rule before running candidate methods.
  4. Build privileged positive controls and strong text-only or surface-mining baselines. Include corrupted explanations, shuffled labels, random directions and no-op interventions.
  5. Run a small control pilot to estimate variance; choose sample size using the desired effect and independent clusters. Keep a fresh private variant for final evaluation.

Required controls

Match text access, number of decoder calls and explanation/query budgets. Surface baselines should include subword cues. A corrupted description should score worse on the concept it was supposed to explain.

Measurements

Primary prediction loss or calibrated accuracy, premium over the strongest baseline, family/model-cluster intervals, control separation and total inference cost.

Deliverable

benchmark_contract.md, versioned splits, positive/null control outputs and one comparison-ready evaluator.

Decision / stopping rule

If the privileged control cannot beat the baseline at the declared effect, or a corrupted method passes, repair the task/metric before comparing interpretation methods.

First action: Choose one retained organism where an internally informed reader has an explicitly testable advantage under the proposed access rules.
T04 · P1Compare automated explanation methods

Does an explanation predict new activations and intervention effects better than matched baselines?

Feature names and detection scores can look good even when the feature produces repetitive or semantically destructive outputs.

Prerequisites: T03, T06
Compute: Bounded explanation and decoder inference

Procedure

  1. Sample a stratified pilot of frequent/rare, lexical/paraphrase-stable and steerable/non-steerable features, plus known negatives. Treat 100–200 features as a pilot suggestion until power is measured.
  2. Implement four arms: activation-example descriptions, intervention/output descriptions, a combined method, and an adaptive hypothesis-testing agent. Reuse existing flipbook machinery.
  3. Give the adaptive arm the same total query/cost budget as a nonadaptive search arm. Record all examples shown to explainers and keep evaluation examples unseen.
  4. Require a structured prediction, applicable domain, uncertainty and anticipated collateral effects from every explanation.
  5. Evaluate once on the frozen split, inspect failures with blinded annotation, then confirm any useful advantage on another trained substrate.

Required controls

Text-only predictors, random/PCA directions, corrupted descriptions, shuffled feature-description assignments and repeated seeds. Select features without inspecting their final test score.

Measurements

Prediction loss, confident false claims, coverage, intervention success, collateral changes, fluency and cost. Report these separately rather than hiding tradeoffs in one score.

Deliverable

method_comparison.csv, all explanation/prediction records, uncertainty estimates and a short failure atlas.

Decision / stopping rule

A gain must exceed the minimum useful effect and survive independent replication. A fluent description or a point estimate without adequate power is exploratory.

First action: Implement the activation-only and output-only baselines on a small frozen feature list before adding an adaptive agent.
T05 · P1Extend the multilingual pivot experiment

What information changes when a French sentence vector is decoded directly into Chinese or passed through English text or an English-reference vector map?

The completed 24-item pilot demonstrates direct translation, but route agreement and cosine do not identify the internal language of a vector. A greedy omission disappeared with beam search.

Prerequisites: None — can start now
Compute: Translation curation + short decoding; CPU maps later

Procedure

  1. Review the saved greedy and beam-5 generations. Annotate paraphrase differences separately from unsupported details, omissions, role changes and corrupted tokens.
  2. Construct an independently checked parallel dataset with contextualized gender, formal address, number, aspect, negation and role contrasts. Resolve singular/plural vous explicitly and allow valid target-language alternatives.
  3. Keep direct decoding and French/English/Chinese text pivots. Add another non-English pivot before treating English as special. Freeze decoding settings and report sensitivity separately.
  4. On a sufficiently large disjoint training set, fit identity, mean-offset, orthogonal Procrustes and regularized linear French-to-English-reference vector maps. Compare a French-to-Chinese-reference map too.
  5. Tune only on validation; hold out meanings and templates. Decode mapped vectors into Chinese and score semantics independently of SONAR, alongside geometric alignment.
  6. If pursuing an internal-computation claim, add layerwise language/meaning readouts and controlled interventions. This is a separate follow-up to the behavioral translation test.

Required controls

The identity map is mandatory. An extra same-language round trip controls for repeated compression. A shuffled-pair map and norm-matched perturbation help separate learned alignment from arbitrary vector damage.

Measurements

Meaning preservation by phenomenon, unsupported specificity, omissions, corruption, route consistency, paired-versus-contrast geometry and uncertainty clustered by contrast family.

Deliverable

validated_parallel_items.jsonl, annotation rubric, route_comparison.csv, map fit/validation records and held-out generations.

Decision / stopping rule

Do not fit a 1024-D transform on 24 examples and call it generalization. Do not identify English as an internal pivot from cosine or language identification alone.

First action: Inspect the 24-item pilot, especially tu/vous, sister-age specification, bicycle specificity and the Chinese apple-token corruption.
T06 · P1Calibrate SAE inference and compare recipes

Does a faithful Matryoshka recipe yield more useful features than plain BatchTopK at comparable cost and sparsity?

The local BatchTopK evaluation uses batch selection, so an example can change activations when its batch companions change. The old Matryoshka trial also differs from the reference training recipe.

Prerequisites: T01
Compute: CPU calibration + a bounded GPU training comparison

Procedure

  1. Preserve the existing code/results and add a calibrated fixed inference threshold using a separate calibration split.
  2. Check the same examples alone and with different batch companions. Report held-out realized L0, its distribution, FVU, feature frequencies and dead features.
  3. Audit the Matryoshka auxiliary loss and relative weighting of prefix losses. Document every intentional deviation from the recipe.
  4. At one retained width and sparsity, compare plain BatchTopK and Matryoshka+BatchTopK using the same data, normalization and at least two seeds. Record training and inference cost.
  5. Run the T04 evaluation plus paraphrase consistency, hierarchy/absorption controls and semantic intervention checks. Use retained seed pairs for an ensemble pilot only after this baseline is trustworthy.

Required controls

Batch-companion invariance, held-out calibration, matched realized sparsity and reconstruction, seed variation and random/PCA directions. Match total width and active-feature budget for ensembles.

Measurements

FVU and L0 distributions as diagnostics; feature-prediction and causal semantic quality as the deciding outcomes, with cost and uncertainty.

Deliverable

versioned inference calibration, recipe audit, two-seed comparator artifacts and a feature-quality table.

Decision / stopping rule

Do not expand width because an automated naming score improves. Escalate only after a held-out feature-quality gain or an informative diagnosis of the current failure.

First action: Run the retained SAE on identical examples with different batch companions to record the current failure before changing inference.
T07 · P2Repair binding measurement before training

Which reader can recover which role information under real vocabulary and construction transfer?

An unconditional global direction, an entity-conditioned reader and a decoder parser have different information and computational access. Their results must not be collapsed into a single binding verdict.

Prerequisites: T02
Compute: Cached embeddings + calibrated probe fits

Procedure

  1. Recover the earlier conditional-reader and decoder-parser code and evaluate both on identical items and splits alongside H1 readers.
  2. Create additive and structured role/filler rotation or tensor-product positive controls under matched distractors and normalization.
  3. Calibrate power separately for each reader and each encoder. Rotate focal identities and balance labels without alphabetical shortcuts.
  4. Compare unconditioned, query-conditioned, nonlinear and decode-then-parse readers with equivalent target definitions.
  5. Only when the assay passes, compare reconstruction and relational objectives at matched data/capacity. Verify a candidate teacher before retrying distillation.

Required controls

Lexical and construction holdout, label permutation, simple surface readers, multiple planted amplitudes, and positive controls suited to the hypothesized code.

Measurements

Transfer accuracy/AUC, calibration, minimum detectable effect, sample size and confidence intervals per reader/encoder; reconstruction cost of any objective intervention.

Deliverable

reader_task_matrix.csv, structured-control generator, power curves and a narrowly scoped binding conclusion.

Decision / stopping rule

If a positive-control gate fails, report instrument failure. It cannot support an absence claim. A positive decoder result is not evidence of a universal linear role axis.

First action: Make the exact task-definition matrix before fitting any new probe.
T08 · P2Separate decoder failures from monitor failures

Can a monitor abstain on semantic errors at useful coverage, and how much does decoder training change the answer?

Round-trip cosine detected some gross perturbations, but a threshold fitted on the same sweep has not established transfer to subtle, fluent meaning changes.

Prerequisites: T01, T02
Compute: Small decoder sweeps + held-out calibration

Procedure

  1. Recover H2 operators and judged examples. Verify that a candidate noise-robust decoder checkpoint is actually accessible and compatible before scheduling its comparison.
  2. Prepare clean, grossly perturbed and near-manifold meaning-change sets; include natural text and role/negation edits with full proposition annotations.
  3. Compare base and robust decoders at matched decoding settings. Include no-op, random norm-matched and mean/shuffled-latent controls to reveal decoder-prior effects.
  4. Fit abstention rules on calibration data only. Evaluate risk versus coverage on held-out examples and shifted domains.
  5. Audit unsupported details and off-target propositions, rather than checking only the attribute deliberately edited.

Required controls

Independent semantic labels, separate calibration/test data, latent-ablated decoder priors and same-dose random directions. Any guarantee requires its actual statistical assumptions.

Measurements

Semantic error among accepted outputs, coverage, false rejection of valid paraphrases, intervention success, collateral changes and confidence intervals.

Deliverable

decoder-by-perturbation-by-monitor table, risk/coverage curves and inspectable accepted-error examples.

Decision / stopping rule

A cosine threshold is a diagnostic, not a universal safety certificate. If the robust decoder is unavailable, retain that comparison as an explicit missing arm.

First action: Split the saved H2 examples by independent source before fitting a new threshold.
T09 · P2Evaluate an accessible second autoencoder

Do the findings and interpretation methods survive on an independently built, usable sentence autoencoder?

A second substrate can challenge SONAR-specific findings. Public weights alone do not establish correct installation, good reconstruction or multilingual support.

Prerequisites: None — can start now
Compute: Model download + reference checks + small GPU pilot

Procedure

  1. Start with QwenAR for reconstruction; evaluate LatentSeal next for its smaller bottleneck and robustness objective. Keep environments separate from the retained SONAR stack.
  2. Pin code and checkpoints and reproduce each release’s own reference checks. For QwenAR, honor its Transformers version requirement and diagnose the checkpoint before fresh sentences.
  3. Run the same unseen English length/domain, role, negation, noise and editing examples through encode-to-one-vector-to-decode. Check that original text never bypasses the vector.
  4. Measure latency, peak memory, reconstruction and semantic retention. Evaluate each model at its own supported dimensions; do not reuse SONAR SAE coordinates.
  5. Choose one competent model for T04 replication. Treat multilingual decoding as a separately gated capability rather than assuming it from a multilingual base model.

Required controls

Release-reference reproduction, sentence-alone versus batch checks, fresh text, shuffled/no-op latents and a shared semantic evaluation set.

Measurements

Exact reconstruction and semantic fidelity by length/domain, corruption, intervention collateral effects, inference cost and memory.

Deliverable

candidate_capabilities.csv, reference check logs, generated examples and a justified substrate choice.

Decision / stopping rule

Exclude a broken installation from scientific comparisons. If reconstruction is inadequate, it can be a failure case but not a fair positive-quality challenger.

First action: Inspect QwenAR’s released evaluation examples and environment requirements; reproduce those before benchmarking new text.
T10 · P2Run only claim-changing follow-up controls

Which remaining uncertainty could change a central paper claim with the least new work?

The campaign flagged many possible follow-ups. Their existence does not justify another broad campaign or a list of disconnected measurements.

Prerequisites: T02
Compute: Mostly cached analysis; targeted decoding when needed

Procedure

  1. Rank the remaining claims by how much a plausible counterresult would change the paper and by the smallest experiment that resolves it.
  2. For operator algebra, test held-out constructions, composition order, inverses and collateral changes at the same dose against random directions.
  3. For capacity, freeze the distortion definition and decoder/inverter; use new length curves and handle censored knees explicitly. Keep the historical estimate post-hoc.
  4. For geometry, measure anisotropy and its expected similarity floor directly. Avoid treating an imported theoretical floor as an empirical mechanism.
  5. For dictionaries, reuse matched-frequency/topic subsets and retained seed pairs. Compare stable subspaces as well as one-to-one feature matches.

Required controls

Matched subsets and denominators, fixed decoders, held-out data, null directions and uncertainty over the actual independent units.

Measurements

A per-claim primary outcome and decision rule; report negative controls and effects on the original headline.

Deliverable

ranked_followups.md and one compact result per selected crux; explicitly defer the rest.

Decision / stopping rule

Do not run a measurement unless its plausible outcomes would alter a claim, method choice or deployment decision.

First action: Select one operator, capacity, geometry or dictionary crux after T02; write its decision rule before executing.
T11 · P2Finish a paper with defensible scope

What is the strongest coherent contribution supported by the reconciled evidence?

The draft has substantial body text but missing publication figures and related work. Its finished-section labels should not conceal unresolved scientific scope.

Prerequisites: T02
Compute: Writing + reproducible plotting

Procedure

  1. Rewrite the abstract around the final scoped contributions and their limitations. Keep SONAR observations separate from ladder-only results.
  2. Add primary related work on sentence-role probing, ROLE, conditional probing, SONAR/LCM and automated interpretation before asserting novelty.
  3. Generate figures from provenance-linked data: binding plus power, capacity with scope/censoring, editing doses, auditing risk/coverage, dictionary stability and benchmark method comparison as results become available.
  4. Include instrument failures, retractions and negative results in the evidence table. Label human-proxy judgments as proxies.
  5. Perform a cross-document pass so paper, plan, board and release describe the same statuses.

Required controls

Trace every plotted number to a file and computation; check that figure captions state the relevant split, reader, substrate and uncertainty.

Measurements

Citation coverage, reproducible figures, resolved contradictions and explicit missing experiments; not the heuristic ledger posterior.

Deliverable

updated paper source, figure scripts/artifacts and a claim-to-evidence appendix.

Decision / stopping rule

Writing can proceed after claim reconciliation. New experimental sections remain provisional until their own task gates pass.

First action: Draft the related-work outline and a six-figure evidence map using only results that survive T02.
T12 · P2Package a reproducible release candidate

Can someone reproduce the benchmark contract and its controls without the original workspace?

The archive is rich but has hard-coded paths, mixed historical snapshots and packages that have not been validated as a public release.

Prerequisites: T03, T11
Compute: Packaging + fresh-checkout smoke

Procedure

  1. Create a minimal release layout with data provenance, model requirements, licenses, pinned environments and accessible small examples.
  2. Replace workspace-specific paths with documented configuration. Package control data and score definitions alongside the implementation.
  3. Provide a CPU instrument smoke and a small GPU model smoke, including one known-valid result and one intentional instrument failure.
  4. Validate the candidate in an isolated fresh checkout and record time, dependencies and downloaded asset sizes.
  5. Prepare release notes stating supported substrates, known limitations and reproduction commands. Keep public publishing as a separate explicitly executed action.

Required controls

Fresh checkout, no implicit cached files, reproducible expected outputs and an instrument-failure example that fails clearly.

Measurements

Smoke reproducibility, missing undeclared dependencies, portability and agreement with the documented benchmark contract.

Deliverable

release candidate directory, licenses/provenance, reproducibility log and release notes.

Decision / stopping rule

Packaging success does not promote the underlying scientific claims. Do not label the artifact publicly released until it actually is.

First action: List the smallest files needed to run T03’s positive and null controls without the rest of the campaign archive.

Evidence, priorities and reading

The source synthesis below preserves completed work, model choices, the 4l inventory summary and reading linked to decisions. The expanded protocols above add implementation detail.

Automated TAE interpretation: things to do

Consolidated 9 September 2026. This combines the four project pages you confirmed, the research/status audit, the SAE review, the 4l inventory, and your multilingual-embedding idea. It is a prioritized working plan, not a claim that the proposed experiments have run. Completed during this synthesis: a 24-item SONAR multilingual pilot, run with greedy decoding and beam size 5 on 4l.

The direction I recommend

Build a benchmark that rewards explanations for predicting unseen behavior, then use it to compare interpretation methods on SONAR and one accessible alternative. Reuse the existing archive. The most useful new exploratory branch is multilingual information retention: what survives direct translation, what an explicit English pivot removes, and whether a language-specific transformation is useful at all.

Keep SONAR as the multilingual baseline. Try QwenAR next for reconstruction and out-of-family interpretability, and LatentSeal for a different bottleneck size and noise-robustness objective. Neither is presently established here as a multilingual SONAR replacement. No omniSONAR migration or reproduction project is on the immediate list.

If choosing just three next work sessions: validate the benchmark's positive and negative controls; examine the multilingual pilot and design its held-out extension; fix SAE inference calibration before comparing existing dictionaries.

What the four pages contribute—and what needs updating

Source Useful material to carry forward How this plan treats it
Research roundup, bibliography source New TAEs, interpretation methods, inversion and latent-generation work Prioritize actual sentence-vector encoders with accessible decoders. Distinguish software releases, papers, and latent-sequence models. Preserve the roundup's 7 September cutoff.
Campaign100 paper A coherent account of binding, capacity, editing, fabrication and dictionaries Complete figures and related work after narrowing claims. Its “DONE” sections are draft-status labels, not universal scientific conclusions.
Broader literature synthesis and its strands Prior art, stronger controls, decoder confounds, competing mechanisms Convert suggested tests into tasks. Do not inherit blanket “nobody has done X,” universality probabilities, or impossibility claims without checking primary evidence and assumptions.
Experiment board, board data All 100 rows, per-claim follow-ups, release and writing backlog Retain outstanding dependencies; do not treat all 63 follow-ups as funded priorities or automatically promote claims based on ledger scores.

Supplementary evidence: status audit, SAE review, remote file inventory, H1–H3 consolidation, FABLE extension debrief, and TRACE/NIGHT8 report. Where reports disagree, trace the exact assay and result file; the newest narrative is not automatically the most accurate.

What is already done

  • Campaign100 has 100 recorded experiment rows and H1–H3 consolidation. Some rows ended in instrument failure or a blocked scientific question; “ran” does not mean “answered.” The paper remains unfinished, with no completed publication-figure set and an empty related-work section in the inspected draft.
  • TRACE's corrected strongest-baseline comparison has not established a robust automated-discovery advantage. One corrected full-baseline premium is 0.00 with a wide interval, approximately [−0.167, 0.222]. The earlier apparent advantage exploited unequal access to raw input and labels recoverable from surface text. This is a useful benchmark lesson, not evidence that interpretation can never help.
  • H1's English cross-vocabulary/construction results are near chance: linear AUC 0.5092, MLP 0.4952. Its additive positive control supports the linear assay at the stated effect size; the MLP fails the same-amplitude gate. Japanese positives did not survive fresh-lexicon testing. These are reader- and split-specific results.
  • Query-conditioned readers and decode-then-parse already have positive results in the earlier FABLE program. Recover their exact settings and reconcile them with H1 before describing conditional probing as new work or claiming role information is absent from z. This corrects the initial status note's presentation of that proposed next experiment.
  • H2 already implements monitored latent rewriting. Natural-text results include negation success around 0.70 and tense around 0.85 in selected dose windows. Above/below is strongly context-dependent. A round-trip cosine detector was calibrated on its own sweep; independent calibration and evaluation are still needed.
  • H3 finds highly readable order within fixed token multisets but weak transfer across pairs. This supports content-dependent readout structure; it does not uniquely identify the causal mechanism. The older binding_death localization attempts had instrument failures and should not be cited as successful causal localization.
  • Matryoshka was already tried locally. Its saved full-dictionary FVU is approximately 0.3104 versus 0.311 for plain BatchTopK in the running report. That exploratory reconstruction tie does not settle feature quality. W40 already includes widths and seed pairs; matched-topic null outputs also exist despite older “owed” notes.

Priority 0: make the existing evidence usable

  • [ ] T01 — Make a reproducible asset/claim manifest. Start with the existing 4l inventory. Record checkpoint type, exact path, model/config revision, normalization, dataset/split, script version and result provenance. Load one representative TRACE checkpoint, one c_pool SAE and one whole-z SAE; verify dimensions and a known reference output. Investigate missing smaller whole-z paths before assuming retraining is necessary. Do not mix c_pool with pooled z.

Deliverable: a manifest and a small reproducible load report. Done when: every asset selected for the first comparison loads and reproduces its reference behavior, or is explicitly excluded. CPU/cached work first; short GPU checks only where needed.

  • [ ] T02 — Reconcile claims, paper and board. Map FABLE conditional-reader positives against Campaign100/H1 negatives by task, query information, construction, vocabulary and power gate. Correct stale claims about matched-topic controls, the 16% atom-energy “ceiling,” high-rank estimates and failed localization. Split fabrication into gross perturbation detection versus subtle meaning changes. Keep the approximately 460-bit capacity figure explicitly post-hoc and substrate-specific.

Deliverable: one claim table with supported scope, counterevidence, result paths and outstanding checks. Recompute the historical ledger only after correcting its inputs; label its values as heuristic evidence aggregation, not calibrated scientific probabilities. Done when: abstract, limitations and board agree on each claim's scope. Formal tier promotion remains an evidence decision, not an automatic bookkeeping task.

Priority 1: the automated-interpretation benchmark

  • [ ] T03 — Freeze a small benchmark contract. Use TRACE for trained organisms with testable internal claims; GAUGE for synthetic instrument checks; SIEVE for encoder capabilities. Reuse TAE-Bench's stimulus and measurement code without conflating these three purposes.

Define three separate outcomes: prediction of unseen feature activation; prediction of a latent intervention's effect; and practical improvement from latent access over matched text access. Freeze what each method can see, its query/compute budget, test splits, primary metric and minimum useful effect before comparison.

Require a privileged positive control that can actually win, alongside shuffled/corrupted explanations, random directions, no-op interventions and strong surface baselines. Hold out propositions, templates and vocabulary; cluster uncertainty by independent family/model rather than treating paraphrases as independent. Use private organism variants for final testing.

Deliverable: a specification, a compact frozen dataset and a control report. Done when: the positive control clears the prespecified effect threshold and known-bad methods fail for the intended reason. If it cannot distinguish these, fix the benchmark before scaling methods. Depends on T01.

  • [ ] T04 — Compare four explanation methods at matched budget. Start with activation-example explanations, intervention/output-based explanations, a combined method, and an adaptive agent that tests competing hypotheses. Give the adaptive method an equal-cost nonadaptive comparator. Reuse the flipbook code where appropriate.

A pilot of 100–200 stratified features is a planning suggestion, not a power calculation. Include frequent and rare features, lexical and paraphrase-stable ones, non-steerable features and corrupted controls. Require every explanation to specify its predicted behavior, applicability and uncertainty. Report held-out predictive loss, confident errors, coverage, intervention success, collateral changes and cost separately. A readable description or fluent decode alone is not success.

Deliverable: one comparison table with uncertainty, failure examples and a small blinded meaning/fluency audit. Done when: the experiment is powered for the chosen effect and replicated on a second trained substrate, or yields an appropriately scoped negative. Depends on T03 and, for SAE comparisons, T06. Do not expand dictionary width merely because the explanation detector score rises.

Priority 1: multilingual embeddings and your English-pivot idea

  • [ ] T05 — Test shared meaning versus an English bottleneck. The initial pilot is complete using the existing SONAR encoder and decoder on 4l; see results and interpretation and pilot code. Direct French→Chinese preserves simple role and number contrasts without an explicit English text intermediate. Adding an English pivot changes outputs, but this small sample does not establish a better route or an internal English bottleneck. A greedy omission disappeared under beam search. The held-out extension below remains to do.

Compare D_zh(E_fr(x)) with D_zh(E_en(D_en(E_fr(x)))). Include an extra French round trip and an extra Chinese round trip, because any decode/re-encode step can lose information. Also compare embeddings of independently supplied equivalent French, English and Chinese texts, with same-topic wrong-meaning minimal pairs.

A successful direct French→Chinese decode shows that a separate explicit English-text step is unnecessary. It does not show the encoder's internal computation contains no English-associated structure. Conversely, matching an English translation's embedding does not make the vector literally an English encoding. Geometry alone cannot establish a unique internal representational language.

Next extension: collect independently checked translations and targeted information-loss contrasts: grammatical gender where English leaves it unspecified, formal/informal address with singular/plural context resolved, tense/aspect, negation and who-did-what-to-whom. Score meaning with bilingual annotation, allowing valid alternatives. Treat a backtranslation or SONAR cosine as supporting diagnostics, not independent ground truth.

Then test a vector translator: on a sufficiently large disjoint parallel training set, compare the identity map, mean language offset, orthogonal Procrustes and regularized linear map from French vectors to English-reference vectors. Tune only on validation; split by meaning/template, not paraphrase. Evaluate held-out Chinese decoding and retained distinctions as well as geometric fit. Compare French→Chinese-reference and French→English-reference alignment to test whether English is privileged. An improvement would establish utility of that fitted transformation, not an internal English pivot. Do not fit a 1024-dimensional map on this tiny pilot and call it generalization.

Deliverable: a route-by-phenomenon report, full generated texts, embeddings and a held-out mapping experiment. Done when: the routes' information loss is quantified independently of the encoder's own similarity metric. The pilot can finish now; the larger study depends on validated translations and T03-style controls.

Priority 1: SAE methods with a trustworthy comparison

  • [ ] T06 — Repair inference calibration, then compare BatchTopK and Matryoshka. The local BatchTopK encoder applies batch selection at evaluation, making a sentence's activations depend on its batch companions. Implement a calibrated fixed inference threshold and verify held-out batch invariance, realized L0 distribution, FVU and dead features. Keep the original outputs/version for provenance.

Use plain BatchTopK as baseline and a faithful Matryoshka+BatchTopK recipe as challenger. Audit AuxK and the relative weighting of prefix losses; the local Matryoshka loop omits AuxK. Match data, normalization, training budget, sparsity and at least two seeds at one existing width. Evaluate feature prediction, paraphrase consistency, absorption/hierarchy controls and causal edits with content preservation, not reconstruction alone.

Deliverable: an audited inference implementation and one controlled comparison linked to T04. Done when: any claimed feature-quality gain survives held-out evaluation and seed uncertainty. If resources allow, test retained-seed ensembles at matched total width and active-feature budget before training bigger dictionaries. Matryoshka+JumpReLU, Group Bias Adaptation and Matching Pursuit are later challengers, not prerequisites. Detailed sources and local code audit.

Priority 2: the most informative scientific extensions

  • [ ] T07 — Repair the binding assay, then test a cause. Reproduce the existing conditional readers and decoder parser on identical items. Add calibrated role/filler-rotation or tensor-product positive controls alongside additive shifts. Verify power per reader and per encoder, including the MLP. Rotate focal identities and hold out constructions and vocabulary. Only after this works, compare reconstruction and relational objectives at matched capacity/data. Row 009 needs a verified binding teacher; row 051 needs per-embedder calibration. A failed gate ends the inference, not the scientific question. Depends on T02.

  • [ ] T08 — Test decoder dependence and selective auditing. Compare the base SONAR decoder with an accessible noise-robust decoder checkpoint if its actual availability/compatibility is verified. Reuse H2 edits; include near-manifold role/negation changes, natural text, random/no-op edits, and mean/shuffled latent baselines that expose decoder priors. Calibrate abstention on a separate set and report risk versus coverage on held-out domains. Check full proposition preservation. A cosine threshold is not a universal safety certificate. Deliverable: decoder × perturbation × monitor table, including failures.

  • [ ] T09 — Try one accessible alternative TAE. QwenAR is the first reconstruction candidate; LatentSeal is the first robustness candidate. First reproduce each release's installation/reference checks in a separate environment. Then run identical fresh English reconstruction, binding, noise and editing tests, with length and domain stratification. Include model size, latency and inference memory; compare usable quality rather than assuming equal vector dimension means equal capacity. Evaluate the text bottleneck alone, ensuring no original-text bypass. Deliverable: a small encoder/decoder capability table and one chosen second substrate for T04. Multilingual support must pass a separate test.

  • [ ] T10 — Finish the strongest remaining controls only if they affect a claim. For operator algebra: held-out constructions, composition order, inverse operations, random norm-matched directions and collateral changes. For capacity: fixed decoder/inverter, explicit distortion definition and fresh length curves; test the post-hoc ceiling on new data. For geometry: directly measure anisotropy and compare its predicted similarity floor with observed “unbinding,” rather than importing a universal VSA floor. For dictionaries: matched-frequency/topic comparisons and seed/subspace stability using the existing W40 outputs. Rank these by which paper claim survives or falls on their outcome.

Priority 2: write and package what is defensible

  • [ ] T11 — Finish the paper from the reconciled claims. Add related work before novelty language: earlier sentence-role probing, ROLE, conditional probing, SONAR/LCM and current automated-interpretation evaluations. Make figures directly from provenance-linked data: binding plus power controls; capacity with censoring/scope; natural editing doses; auditing risk/coverage; dictionary seed stability; and benchmark comparison. Preserve negative and instrument-failure rows. A ledger visualization is supplementary evidence accounting, not the main result. Depends on T02; new-results sections depend on the relevant tests.

  • [ ] T12 — Prepare a reproducible local release candidate. Document datasets, model requirements, licenses, pinned environments and CPU/small-GPU smoke commands. Validate in an isolated fresh checkout. Fix package-path assumptions and stale control claims. Include examples of both a valid result and an instrument failure. Public publishing is a separate step; this plan does not send messages, publish, or authorize a paid compute campaign.

Human calibration from row 060 remains real unfinished work: its existing panel was an LLM proxy. Use human annotation first where it resolves the multilingual and meaning-preservation labels, rather than adding a disconnected general similarity study.

What 4l already gives us

Paths below are relative to /workspace/HOME/guest/ on 4l; sizes are approximate audit snapshots. File existence was checked, not universal checkpoint loadability.

Collection Retained assets Immediate use
night8/ (~8 GB) 10 checkpoint files and corrected benchmark boards T01, T03, T04
campaign100/ (~16 GB), consolidation/ (~94 MB) Experiment records, H1–H3 outputs, steering code T02, T07, T08, T11
ladder/ (~13 GB) 28 final checkpoints including BART and objective variants Positive controls, objective comparisons; BART was not previously a strong-quality replacement
binding_death/ (~7.8 GB) 52 cached layer-state files Conditional-reader/assay work after fixing the old instrumentation problems
w40/ (~6 GB) Eight c_pool dictionaries, two 131k whole-z dictionaries T06 and seed/stability controls
statspass/ (~560 MB) Stimuli and cached SONAR embeddings Cheap cached-data reanalysis
parascopes_layerwise/ (~23 GB) Broad SAE/interpretation code and artifacts Reuse the existing explanation workflow

The audit saw four 16 GB A4000s available at the instant checked, not a standing reservation. Begin with cached CPU analysis and short decoding jobs; measure throughput before scheduling training. No inference from an idle snapshot to a promised training completion time is intended.

Model choice and the parked omniSONAR discussion

Candidate Reason to try it Main qualification
SONAR Existing multilingual encoder/decoder and archive; immediate pivot experiment Strong baseline for this program, not guaranteed language-invariant or universally faithful. Official project
QwenAR Public single-vector 1024-D autoencoder, built from 0.6B-class encoder/decoder components English training; the reported 30/32 exact reconstruction result is a tiny probe set. The project requires Transformers ≥5.5 and reports silent degradation under 4.x. Reproduce its reference check before fresh evaluation. Project, weights
LatentSeal Public independent text encode/decode interface, 256-D unit-norm bottleneck; useful robustness comparison Watermarking-oriented and tagged English. Verify text-only quality and supported lengths; multilingual suitability is unestablished here. Model card and API examples, paper

OmniSONAR stays a reading item. No public checkpoint/API was verified in the access audit; an author said on 31 July that release dates—and release itself—could not be promised. Author comment.

For reference, its paper reports a 1.5B text encoder and 1.8B decoder with a 1024-D bottleneck. Headline SONAR→OmniSONAR comparisons include FLORES xsim++ error 15.3→6.1, a selected MTEB average 63.338→74.114, and FLORES translation chrF 52.9→55.4. These establish useful benchmark gains, not improved SAE interpretability or exact English reconstruction. The recipe combines LLM initialization, translation, pooled/contrastive learning, hard negatives and language expansion. Paper.

The earlier $20k–100k successful-text-run estimate, centered roughly on $50k, was a speculative planning range, not Meta's reported bill or a reproduction quote. It assumed 128 GPUs for 125k core steps at 1–5 seconds/step (about 4.4k–22.2k GPU-hours), added the reported expansion stage and uncertain overhead, and used illustrative public A100 rental prices. It excludes original Llama pretraining, data production, speech, the distilled-model suite, research reruns and labor. Timing and hardware assumptions dominate the uncertainty. It is not a reason to rebuild an unavailable model for this project. Pricing examples used, Runpod cost guide.

Reading tied to decisions

Read when Paper Decision it informs
Before freezing T03 CHIVE Counterfactual prediction and strong transcript-only baselines
Before T03/T04 Pitfalls in Evaluating Interpretability Agents Memorization, leakage and private evaluations
Before T04 Output-Centric Feature Descriptions, Automatically Interpreting Millions of Features Practical input/output explanation baselines
Before the adaptive arm Automated Interpretability and Feature Discovery with Agents Competing hypotheses and targeted controls
Before T06 BatchTopK, Matryoshka, Are SAE Benchmarks Reliable? Faithful recipes and metrics that can distinguish methods
Before T07 and related work Conditional Probing, ROLE Reader-relative information and role/filler structure
Before claiming generality SynthSAEBench, InterpBench, Explanation Sanity Checks Ground truth, model variation and corrupted-description controls
Before T08/T09 Large Concept Models, QwenAR and LatentSeal above Decoder training as a confound; accessible alternative substrates
Later CALM, vec2vec, SentenceLens Chunked bottlenecks, cross-space mappings and activation decoding as separate future branches

Do not turn the whole bibliography into prerequisites. CHIVE, Pitfalls and the output-centric paper are enough to begin the interpretation benchmark; BatchTopK, Matryoshka and the reliability audit are enough to begin the SAE comparison.

Deferred until a smaller experiment justifies them

Another broad 100-row campaign; a larger SAE merely to improve description scores; a from-scratch omniSONAR reproduction; sweeping universal claims about pooled representations; and moving wholesale to latent-sequence/diffusion models. The latter are interesting, but change the experimental object. A deterministic embedding of a complete input cannot add unrestricted Shannon information to that input: practical uplift must be defined through constrained access, computational budget, or model-specific internal/causal predictions.