Existing evidence and provenance
This register supports the new manuscript. Reports were read and selected on 11 September 2026; no original experiment was rerun for this project. Source snapshots preserve exact report text, including stronger historical interpretations that the scoped summaries below do not endorse. The benchmark's shared method comparison remains [needs to be run].
Original paths, full-source hashes, snapshot hashes, and any excerpt boundaries are in evidence_sources.json. Downloadable text snapshots make these citations usable independently of the old site's publishing coverage. These are reports, not a substitute for the raw prediction arrays needed for final figure reproduction.
Figure update: Figures 3–4 now use frozen numeric data with source hashes extracted from original JSON artifacts. Figure 4's five swap/retention cell means and primary paired point contrast were independently reaggregated from retained output rows. Saved bootstrap intervals are displayed without recomputing them. No model inference or new experimental run was performed.
S01 — Composer installation on a trained miniature TAE [existing result]
Source: archived panel-response excerpt, original tae-interp/docs/debriefs/PANEL_RESPONSE.md, installation/ablation lane.
Use: §5.1; a candidate constructive case study. B_D_s0 full-composer oracle-output token-F1 0.551; symmetric average 0.366; matched-spectrum random rotation 0.450; rotation removed 0.509. Three pair-seed means on this named model. Order discrimination 0.994 full, 0.962 with rotation removed, and 0.500 for symmetric averaging.
Scope: decoder utility of an installed fitted composer on a small trained organism. Not a SONAR encoder-mechanism demonstration. Pair seeds are not independent model-training seeds. The best tested angle is not proof of a continuous global optimum. Avoid repeating the source's stronger mechanism language.
Figure 3 completed: selected means are computed from the original three pair-seed aggregates, with individual pair seeds plotted. The saved OD metric averages forward and swapped directions on scorable cases; these need not be identical cohorts across arms. No confidence interval is inferred from three pair seeds. [needs to be improved] Full per-item/interval reproduction before submission. [needs to be run] N06a transfer under the shared semantic contract.
S02 — Query-conditioned order readout versus editing [existing result]
Source: E04b report, fresh 160-proposition SONAR experiment, 11 September 2026.
Use: §5.2; default second case study. Reader AUC 0.977–0.980 across construction families. Reader-subspace swap rate 0.140; two structured-control rates 0.124/0.110; full counterfactual difference 0.995. The primary paired comparison chooses the best energy-matched control separately for each proposition, then averages the difference: −0.00104, saved interval [−0.0395, 0.0359]. Selecting a single best arm across the full dataset instead gives +0.015625, saved interval [−0.0213, 0.0523]. The figure update verifies the primary point contrast against retained rows and the scorer; the difference in estimand is now resolved.
Scope: no resolved selective advantage for the tested readout subspace over the tested matched controls. Does not rule out all low-dimensional causal handles or all relational representations. Readout performance and interventions are different endpoints.
Figure 4 completed: swap/retention point estimates verified against retained rows; saved noun-cluster intervals imported. [needs to be improved] Reproduce the bootstrap and extend preservation beyond the retained-noun check. [needs to be run] Shared-protocol comparison in N05/N06b.
S03 — Lexical/local anatomy of trained SONAR [existing result]
Source: E01 report, 8,000 natural sentences / approximately 226,000 tokens, document-disjoint split.
Use: §5.3; motivates residual information as an open question. Pooled-z centered R² 0.233 for F4. Own-token linear contribution approximately 0.370 on content-token states; random-init F4 approximately 0.938 under its own token-state evaluation.
Scope: token-state and pooled-vector numbers have different targets and must not be pooled. Poor prediction is model/data/budget-specific. The sampled bigram estimator recovers only 4% of a generic planted bilinear term, so a small interaction gain is not evidence of small true interactions. The report explicitly gives no identified semantics for the unexplained component.
[needs to be run] N07 characterization of residual lexical and sentence information on controlled held-out combinations. Do not claim words have been fully removed.
S04 — Natural-text steering and dose limits [existing result]
Source: H2 report, consolidation of existing negation/tense/vertical operators.
Use: §4.2 / potential N06a extension. Reported fluent-gated negation success approximately 0.70 at dose 1; tense approximately 0.85 at doses 1.5–2. Higher doses impair fluency. Above/below transfer was approximately 0.04 on the tested natural usages.
Scope: selected-dose, single-system, exploratory judged measurements. These are not independently validated all-content-preserving edit rates. The source's perfect-specificity and deployment claims are not adopted; later audit/transfer work narrowed them. A monitor calibrated on its own sweep does not establish a general guarantee.
[needs to be improved] Recover exact endpoint denominators and source-level uncertainty. [needs to be run] Independent dose calibration, applicability labels, and collateral scoring in the shared evaluation.
S05 — Unbinding and metric dependence [existing result]
Source: U1 report, frozen artificial cross-document chains, 2,048 training / 512 calibration / 1,024 test master chains.
Use: §5.1; reconstruction versus retrieval distinction. At training N2048, centered ridge cosine falls 0.6631 → 0.1341 from depth 2 → 16. At depth 16, raw ridge cosine 0.3414 versus mean baseline 0.3142. Retrieval@128: ridge 0.0295 versus direct-chain 0.2891.
Scope: depth changes content count and token length together. Retrieval uses unrelated-chain distractors, not same-chain position discrimination. Conditional intervals do not represent model-training or lexical-population uncertainty. No information-capacity theorem follows from this fixed estimator's decline.
[needs to be improved] Keep metric definitions attached to plots; use as supporting evidence rather than adding a full capacity section.
S06 — SAE invariant masks: task-dependent specificity [existing result]
Source: S2 controlled-bank report, 288 clusters in 16 lexical blocks, eight historical dictionaries, two inference rules.
Use: §4.2; reason to evaluate decomposition by multiple capabilities. Per-sentence TopK invariant-minus-matched-mask rank-utility differences: role −0.0934, polarity +0.3998, count +0.1584. Threshold-inference signs agree.
Scope: a fixed-bank mask comparison, not an all-method leaderboard, a causal edit, or evidence of monosemantic atoms. Native full-z cosine already solves polarity/count on this bank. Matching and selection confounding remain; harder contrasts and subsequent activity-matching audits narrow broad specificity interpretations. Poor role cosine does not show role information is absent.
[needs to be improved] Reconcile this selected result with later activity-matching corrections before making a new specificity headline. [needs to be run] N05 common dictionary comparison; do not silently combine whole-z with c_pool checkpoints.
S07 — Testing a sentence-broadcast account [existing result]
Source: E03 report, same sentence collection as E01.
Use: §5.3; a tested alternative explanation, not a new headline. Adding a linear contribution from other final token states raises content-token R² approximately 0.474 → 0.509; a tested interaction adds approximately 0.016. A random-init counterpart yields a much larger broadcast gain.
Scope: the particular shared-state account leaves substantial unexplained variance at the tested budget. This does not identify the remaining features or establish that no other contextual model can explain them. E01/E03 share data and are not independent replications.
[needs to be improved] Keep the comparison in the appendix unless it directly explains the final method comparison.
S08 — TRACE's corrected baseline comparison [existing result]
Source: maker-r9 report; access/scoping cross-check from the September audit.
Use: §4.2 and the benchmark access rules. Corrected distill_verify premium 0.000, 95% interval approximately [−0.167, +0.222], on the v2 basis_full+ comparison. The audit traces this to the saved r9_v2 result JSON.
Scope: no statistically resolved advantage in this method/task/budget comparison. Not proof of equivalence or that latent interpretation cannot help. Strong baseline matching corrected an apparent gain from unequal access to input-derived labels. Do not adopt historical claims of universal benchmark novelty, perfect certification, or an airtight evaluation.
[needs to be improved] Recover raw predictions and the exact baseline contract for a reproducibility check. [needs to be run] The new common comparison N05; these scores cannot fill its cells.
What is deliberately not imported
Universal absence of binding; a uniquely identified SONAR rotation mechanism; intrinsic 460-bit capacity; a universal semantic energy fraction; calibrated scientific probabilities from the historical evidence ledger; deployable semantic monitoring guarantees; and a claim that all linear or nonlinear readers fail. These exceed what the selected evidence supports or are not necessary for this paper.
Figure provenance and evidence promotion
An existing report is sufficient to seed prose with a scoped label. A final data figure additionally requires original numeric artifacts, metric/denominator checks, and reproducible aggregation. A new shared-benchmark claim requires the frozen evaluation itself. Keep those three states distinct when updating the manuscript.