# Sample-size and compute audit of the 100 campaign rows

Generated 2026-09-11 from a read-only audit of every row's `PREREG_LITE.md`, `RESULT.md` and `out/` files.
Flags: **ok** = hundreds or more independent items with clustered uncertainty, or an exact count; **thin** = fewer than ~200 items behind the headline, or no clustered uncertainty, or a single trained seed where the claim depends on training; **very thin** = fewer than ~50 items, or a single seed of a trained model with no repeat; **n/a** = instrument, theory or packaging row with no data claim of its own.

| row | flag | items behind the headline | training set | seeds | uncertainty | compute | substrate | why |
|---|---|---|---|---|---|---|---|---|
| 001 | very thin | 500 held-out test sentences (50 test propositions) per pooled cross-construction+lexical cell; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences) | organism: 5e8 target tokens (43,898 steps) OWT DAE + 20% mix of 140k role-swap pairs; probe: 150 train propositions (300 items) per family | 1 training seed (seed 0); battery probe seeds 0,1,2 | 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ladder ~11.6M params (d_model 256, 4+4 layers), 5e8 target tokens, 1 GPU (GPU0 on 4l), 5h max-hours guard, ~6 GPU-h cap incl. battery; GPU type not stated | TAE ladder (trained in-house, rung A) | Single training seed, single lambda, single rung, confounded control (mix-in + InfoNCE); null on transfer rests on one trained model with no repeat. |
| 002 | very thin | 500 held-out test sentences (50 test propositions) per dose; 3 doses; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences); aux-head manipulation check on a proposition-level 10% held-out split of the 140k corpus | organism: 1.7e8 target tokens per dose (1/3 of 001) + 140k role-swap corpus (90/10 split); probe: 150 train propositions per family | 1 training seed per dose (3 doses); battery probe seeds 0,1,2 | 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ~11.6M params, 1.7e8 target tokens per dose, 3 sequential runs on GPU2, ~5-6 GPU-h total; GPU type not stated | TAE ladder (trained in-house, rung A) | Single seed per dose, three-point ladder, budget 1/3 of baseline; lambda=0.01 head collapsed; dose-response null from three unrepeated trained models. |
| 003 | very thin | 500 held-out test sentences (50 test propositions) for the transfer null; manipulation check on 2k held-out transitive sentences incl. 500-item swapped-order subset; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences) | organism: 5e8 target tokens, 20% transitive structured-target mix (140k-pair corpus, 117 content lexemes); probe: 150 train propositions per family | 1 training seed (seed 0); battery probe seeds 0,1,2 | 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds; no CI stated for the 0.998 manipulation-check accuracy | rung-A ~11.6M params, 5e8 target tokens, GPU0 on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A) | Single seed/rung/budget; transfer null from one trained model; manipulation check numbers reported without CI. |
| 004 | very thin | 500 held-out test sentences (50 test propositions) for surface-cross and primary cells; reorder manipulation check on 500 held-out OWT + 500 held-out transitive sentences; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences) | organism: 5e8 target tokens, full word-shuffle corruption, 20% transitive mix; probe: 150 train propositions per family | 1 training seed (seed 0); battery probe seeds 0,1,2 | 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds; no CI on Kendall-tau | rung-A ~11.6M params, 5e8 target tokens, phys GPU2 on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A) | Single seed of a trained model; the 'twist' (surface AUC 0.615 vs baseline 0.706) is a single-run comparison with no training-seed repeat. |
| 005 | very thin | 500 held-out test sentences (50 test propositions) per pooled cell; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences) | organism: 5e8 target tokens, 40% decorrelated transitive mix (every proposition x 4 families x 2 orders); probe: 150 train propositions per family | 1 training seed (seed 0); battery probe seeds 0,1,2 | 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ~11.6M params, 5e8 target tokens, GPU0 on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A) | Null from a single training seed, single rung, no lambda=0 same-mix control; mix fraction differs from 001-003 so cross-arm comparison is caveated. |
| 006 | very thin | 300 novel-filler sentences (350-word reserved noun slice, 8 shared verbs, 4 families) per dose for the headline 0.00/0.05/0.73; 2k held-out same-vocab sentences for the lookup check; 500 test sentences for the battery cell | 1.7e8 target tokens per dose; transitive corpus 252k train / ~28k val rows at every N (N_lexemes in {30,117,1000}) | 1 training seed per dose (3 doses) | none stated for the novel-filler accuracies (no CI); battery cell has 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ~11.6M params, 1.7e8 target tokens x 3 doses sequential, phys GPU2 on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A) | The block's headline positive rests on one seed per dose at 1/3 budget, 3 grid points, and 300 novel-filler sentences with no CI; RESULT itself calls for 3 seeds x finer grid before promotion. Parse rate dipped to 0.87 at N=117. |
| 007 | very thin | 2k held-out transitive items for the slot-specialization matrix; 500 test sentences (50 test propositions) for the battery cell; binding battery agent_patient: 2000 sentences = 200 propositions x 2 orders x 5 families; lexical split 150 train / 50 test propositions; primary cell pools 20 ordered family pairs, each tested on the 50 held-out test propositions (500 sentences) | organism: 5e8 target tokens with 003's structured target, 20% mix; slot probes trained on held-out transitive items (split not stated) | 1 training seed (seed 0); battery probe seeds 0,1,2 | none stated for slot accuracies (delta 0.001-0.03 reported without CI); battery cell has 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ~11.6M params + two-slot attention pooler, 5e8 target tokens, 1 GPU on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A, two-slot variant) | Single training seed of a new architecture; the 'redundant twins' finding is one run with no repeat and no CI on the slot matrix. |
| 008 | thin | 2000 agent_patient stimuli (500 test sentences / 50 test propositions per pooled cell) probed at 17 checkpoints x 3 seeds | 3 organisms x 5e7 target tokens (10% of standard budget), pure OWT DAE, no transitive mix | 3 training seeds (0,1,2) x 17 dense checkpoints; probe seeds 0,1,2 | proposition-cluster bootstrap CI with n_boot 200 (reduced from 1000), probe-seed envelope | rung-A ~11.6M params, 5e7 tokens per seed (~8 min/seed), 3 seeds; CPU probe analysis; GPU type not stated | TAE ladder (trained in-house, rung A) | 3 seeds is the best in the block, but 10% budget means the surface code barely formed (RESULT: probe-power gate effectively binding), n_boot 200, and dips of <=0.007 are within selection noise over 17 checkpoints. |
| 009 | ok | 2000 agent_patient stimuli (500 test sentences / 50 test propositions per pooled cell) + 1200 genitive stimuli as positive control; 20 cells (5 layers x 2 poolings) | none (frozen Qwen2.5-3B-Instruct teacher; Phase 2 distillation not run); probe: 150 train propositions per family | 1 teacher model/checkpoint; probe seeds 0,1,2 | proposition-cluster bootstrap CI, n_boot 200, probe-seed envelope (min-lo/max-hi); gate = max CI-lo over 20 cells | one forward pass of Qwen2.5-3B-Instruct (36 layers, H=2048) on 3200 sentences, wall 180 s on GPU0; GPU type not stated | other (frozen Qwen2.5-3B-Instruct pooled hidden states; probes from the SONAR binding battery) | Frozen-model null on 2000 items with clustered CIs and within-family ceiling as power check; caveat: one teacher model, pooled reps only, n_boot 200, 50-proposition test universe. |
| 010 | very thin | 2k held-out same-vocab sentences + 300 novel-filler sentences for the case-accuracy headline (0.503); 500 test sentences (50 test propositions) for the battery cell | organism: 5e8 target tokens, 20% MT-to-case-target mix reusing 006's N=1000 corpus (252k rows) | 1 training seed (seed 0); battery probe seeds 0,1,2 | none stated for case accuracy (0.503 vs 0.5025 baseline, no CI); battery cell has 95% proposition-cluster bootstrap CI (binding_battery.py default n_boot=1000; count not restated in RESULT), envelope over 3 battery probe seeds | rung-A ~11.6M params, 5e8 target tokens, GPU0 on 4l; GPU type/hours not stated | TAE ladder (trained in-house, rung A) | 'Unlearnable' claimed from a single seed at a single budget/lambda with no CI; RESULT concedes a longer run or higher mix might learn it. |
| 011 | ok | 2000 agent_patient sentences (primary dif cell tested on 50 held-out test propositions / 500 sentences pooled over cross-family pairs) + 1200 genitive sentences; single surface L24n | none (frozen SONAR encoder states); linear probes fit on the 150 train propositions per family | 1 SONAR checkpoint; probe/bootstrap seeds inherited from bd_probe_v3 (count not stated in RESULT) | proposition-cluster grid bootstrap 95% CI (resample count not stated in RESULT), permutation test (200 permutations, p=1.0), planted-signal power curve {2,5,10,20,40}% norm x 3 directions, G-repro reproduction of 0.5084 | CPU only, 339 s wall on box; no GPU | SONAR (frozen) | Certified null: 2000 items, clustered CI, permutation, planted-signal power certification (AUC>0.9 by 10% norm). Caveats: one checkpoint, one surface, 50-proposition test universe; projection variant is unpowered. |
| 012 | ok | 2000 agent_patient sentences (200 propositions x 10 realizations); 384 heads x 5 families; held-out split = half the propositions (100) for the top head | none (frozen SONAR attention weights) | 1 SONAR checkpoint; label-permutation null run with 5 seeds | cluster bootstrap SE via 500 resamples of the 200 propositions -> per-family z; Bonferroni max-statistic over 384 heads (z*=3.83); label-permutation control (0 candidates in 5 seeds); held-out half-split replication | GPU2 (CVD=1) attention capture + CPU analysis, whole chain ~15 s; GPU type not stated | SONAR (frozen) | 2000 items / 200 proposition clusters, Bonferroni over 384 heads, permutation null and held-out replication; caveat: bootstrap only 500 resamples, one checkpoint, English only. |
| 013 | ok | 2000 agent_patient sentences (200 propositions x 5 families x 2 orders), each decoded under baseline + 11 ablation conditions | none (frozen SONAR encoder/decoder; mean-ablation) | 1 SONAR checkpoint; random-10 control x 5 seeds, random-42 control x 2 seeds | proposition-cluster bootstrap 95% CI on the preservation delta vs baseline, n_boot 2000 (src/run_knockout.py default) | GPU inference only (encode + greedy decode of 2000 x 12 conditions), no training; GPU type/time not stated | SONAR (frozen) | Null (+0.001 [-0.006,+0.007]) is tightly bounded by clustered bootstrap on 2000 items with count-matched random controls; caveats: mean-ablation may be too gentle, one checkpoint. |
| 014 | thin | 43 kept swap-pairs per family (active, passive) after equal-piece-count exclusions (130 of 200 pairs excluded per family), x 13 layers x 3 conditions; cleft/objrel/nominal secondary | none (frozen SONAR; activation patching) | 1 SONAR checkpoint; random-span control seeded (1 seed) | pair-level bootstrap 95% CI, n_boot 2000; CIs ~[0.88, 1.0] at flip 0.95; parser gold gate 1.000 | GPU inference (13 layers x ~86 pairs x 3 conditions greedy decodes), no training; GPU type/time not stated | SONAR (frozen) | Only 43 pairs per family (65% of pairs excluded by strict piece-matching), but the effect is 0.95 vs 0.00 with bootstrap CIs, so the cliff shape is robust; layer resolution between L22 and L24 is coarse. |
| 015 | ok | 2000 agent_patient sentences; each of 20 ordered family pairs tested zero-shot on 50 held-out test propositions (500 sentences pooled); 4 surfaces (L8/L16/L22/L24n) | attention pooler + linear head trained per family pair on family F1 & split=train (~300 items) on frozen SONAR token states | 1 SONAR checkpoint; pooler init + resample seeds 0,1,2 | proposition-cluster grid_bootstrap 95% CI (battery n_boot=1000 default), 3-seed envelope, planted-signal power certification (lo 0.90-0.95 at L8/L16/L22; marginal 0.648 at L24n), mean-pool baseline anchor | trained small pooler on cached states; GPU/CPU and time not stated in RESULT | SONAR (frozen states) + small in-house pooler | Certified transfer null on 2000 items with clustered CIs and planted-signal power; caveats: pooler is not focal-conditioned (within-ceiling 0.71-0.82), L24n power marginal, ~300 training items per pair. |
| 016 | ok | 2000 agent_patient sentences; 20 ordered family pairs each tested on 50 held-out test propositions; 3 probe families x {L8,L16,L22,L24n} | kernel (RFF D=2000)/focal MLP (256 hidden)/bilinear (rank 16) probes trained per family pair on ~300 train items of frozen SONAR states | 1 SONAR checkpoint; probe init + resample seeds 0,1,2 | proposition-cluster grid_bootstrap 95% CI, 3-seed envelope, planted-signal power curve (certified at eps 2-8x feature-std / theta up to 1.57 rad); L24n UNPOWERED in every family; permutation not triggered | CPU only (<=12 workers); wall not stated | SONAR (frozen) | Adequate n and clustered CIs; but nulls are certified only down to a coarse sensitivity (plant must be comparable to the surface code) and L24n is uninterpretable; prereg amendment broadened the certification after seeing data. |
| 017 | ok | 2000 agent_patient sentences; 20 ordered family pairs on 50 held-out test propositions; 9 position classes x 8 layers = 72 role cells (+72 surface cells) | none (frozen SONAR); linear probe on class-mean states fit per family pair (~300 items) | 1 SONAR checkpoint; probe seeds 0,1,2 | proposition-cluster grid_bootstrap 95% CI, 3-seed envelope; planted-signal power certified at 3 representative cells only; G-repro anchor 0.508 | CPU only, wall 330 s after a block-rule fix; no GPU | SONAR (frozen) | 2000 items with clustered CIs; the heatmap null (role 0.48-0.55 across 72 cells) is powered only at 3 certified cells and reads a 50-proposition test universe. |
| 018 | thin | per cell: 100 test items (50 test propositions) within one family; 8 families x 6 layers x 2 targets = 96 funcword cells; 3200 items total (2000 agent_patient + 1200 genitive) | none (frozen SONAR); linear probe on 300 train items (150:150) per family | 1 SONAR checkpoint; probe seeds 0,1,2 | proposition-cluster bootstrap n_boot=1000, 3-seed envelope (CIs ~+/-0.07-0.13 per cell); planted-signal power certified at 2 funcword cells (eps=0.15 additive, theta=0.6 rotation); randctl noise floor | CPU only, chain wall ~50 s; no GPU | SONAR (frozen) | Each within-family test cell has only 100 items / 50 propositions (CI half-widths up to 0.13) and 96 cells are searched; the null is powered at 2 cells but per-cell maxes up to 0.69 sit inside the multiplicity envelope. |
| 019 | ok | 2000 agent_patient sentences teacher-forced through the decoder; 20 ordered family pairs on 50 held-out test propositions; 8 layers x 3 spans x 2 targets = 48 cells | none (frozen SONAR decoder); linear probe fit per family pair (~300 items) | 1 SONAR checkpoint; probe seeds 0,1,2 | proposition-cluster grid_bootstrap 95% CI, 3-seed envelope; planted-signal power certified at 2 decoder cells; within-family ceiling; behavioral anchor 0.981 | GPU teacher-forced capture of 2000 sentences (fp32, length-bucketed) + CPU probes; GPU type/time not stated | SONAR (frozen) | 2000 items, clustered CIs, power certs; the headline caveat is design (teacher forcing puts order trivially in the input), not n. |
| 020 | ok | 2000 agent_patient sentences teacher-forced (per-class token accuracy); free-running decode ablations on a ~500-sentence subsample (prereg) for role-preservation/chrF | none (frozen SONAR decoder; cross-attention hook ablation) | 1 SONAR checkpoint | none stated (no CI on ablation accuracies); constancy claim is an exact architectural check (per-step std/norm = 0.000 at all 24 layers, cross_src_len=1) | GPU teacher-forced + greedy decode passes over 2000 sentences; GPU type/time not stated | SONAR (frozen) | Headline is a deterministic architectural fact (length-1 cross-attention source) verified on all 2000 items; the secondary ablation effect sizes (e.g. 0.980 -> 0.856) carry no CI. |
| 021 | thin | 1500 clean sentences (500 per SPM-length bin 8-15/16-25/26-40); 54 greedy fabrications (61 nucleus); cosine-bin cells 10/30/124/1336; judge reliability n=150; hand calibration n=40 | none (frozen SONAR round-trip; single LLM judge codex gpt-5.6-sol) | 1 (nucleus sampling seed 21); one judge pass + one V2 re-judge | no CI on the 3.6% rate; judge agreement only (V1 vs V2 binary kappa 0.885 on 150; hand-vs-judge 38/40); RESULT notes bin cells of 9-39 fabrications are 'indicative, not tight' | GPU0 round-trip decode 62 s; judging local codex; GPU type not stated | SONAR (frozen) | 1500 items give a tight binomial rate, but the taxonomy shares (72% entity), per-bin rates, and gate-blindness (35/54) rest on 54 positive events labeled by one LLM judge with no CI reported. |
| 022 | thin | 1500 sentences with 54 fabrication positives; span stats on 38 substituted sentences / 159 tokens; detector test folds have 22 (all) / 14 (gate-blind) positives | logistic detector on 60% stratified split of 1500 (4-5 features); labels inherited from 021 | 3 detector split seeds; 1 SONAR checkpoint | paired per-sentence bootstrap 95% CI on entropy diff (+1.40 [1.00,1.83]; n=2000 resamples per src); detector AUC mean +/- sd over 3 seeds (sd up to 0.054); Cohen's d | GPU0 greedy decode-with-scores 131 s + CPU analysis; GPU type not stated | SONAR (frozen) | Positive effect (d 1.20) rests on 38 fabricated sentences; AUC estimates come from folds with 14-22 positives and no CI beyond 3-seed sd; RESULT itself flags 'small positive n'. |
| 023 | thin | 400 sentences x 19 conditions = 7600 greedy decodes for metrics; judged headline (band existence) on a fixed 100-sentence subsample per condition (1900 labels); altered outputs N=168 for pred (d); hand check 20 | none (frozen SONAR; in-house SAE ckpt defines the off-manifold direction) | 1 (direction seeds fixed per sentence); single judge pass + V2 re-judge on 150 | none stated (no CI); RESULT notes ~+/-0.05 binomial SE at n=100 per cell; judge kappa 0.581 (moderate) on the altered/garbled split | box GPU decode of 7600 sentences + re-encode; local codex judging; GPU type/time not stated | SONAR (frozen) | The 'no fail-open band' null is claimed from 100 judged sentences per (direction, dose) cell with no CI, a moderate-kappa judge, and greedy decode only; the qualitative shape is consistent across 18 cells. |
| 024 | very thin | 127 paired fabricated tokens (123 entity, 4 number) from 38 fabricated sentences; 695 control tokens from 310 faithful-paraphrase sentences; first-token minimal pairs n=48 | none (frozen SONAR decoder; reads-ablated prior) | 1 SONAR checkpoint | sentence-cluster bootstrap 95% CI (n=5000 resamples per src/analyze_prior.py): fab +4.25 [2.96,5.42]; fab-control +1.54 [0.19,2.79] | GPU0 teacher-forced pass, 17 s decode; GPU type not stated | SONAR (frozen) | The load-bearing specificity test rests on 38 fabricated sentences (127 tokens); the fab-control CI lower bound (+0.19) is close to zero and the number-change arm has n=4 tokens; the non-specific 'prior prefers emitted token' effect is stronger than the fabrication-specific one. |
| 025 | ok | 200 base sentences (stratified over 3 length bins); 200 x 20 neighbours x 6 alpha = 24,000 chord decodes + 200 dose-0 + 600 gradient records; 1275 unique reconstructions judged; headline miss rate = 181/200 at T=0.95 | none (frozen SONAR; KNN interpolation, no optimizer trained) | 1 (seed 20250802); single judge pass + V2 re-judge on 150; hand check 20 | none stated (no CI on the 0.905 miss rate); judge V1/V2 kappa 0.838; hand-vs-codex 16/20; G1-G3 sanity gates | box GPU: 24,000+ greedy decodes each re-encoded twice; local codex judging; GPU type/time not stated | SONAR (frozen) | 200 sentences and no CI, but the effect (181/200 accepted at every threshold, mean chrF ~10) is far from any decision boundary; the real caveat is worst-case-per-sentence selection, not sample size. |
| 026 | very thin | semantic rates judged on a fixed 300-sentence subsample (100 per bin) per strategy; fidelity metrics (exact/chrF/cos) on the full 1500; fabrication counts greedy 9 / beam8 11 / beam4 12 / temp1.0 168 of 300 | none (frozen SONAR decoder) | 1 seed (21) per stochastic primary decode; self-BLEU uses 3 seeds on 300 sentences | none stated (no CI); RESULT concedes top-5 strategies are 'not statistically separable' at 9-14 events; hand check 15/15 on beam8 (all one class) | phys GPU2: 7 strategies x 1500 decodes + re-encode; local codex judging; GPU type/time not stated | SONAR (frozen) | The 'beam does not reduce fabrication' null rests on 9 vs 11 vs 12 fabrication events in 300 judged sentences, one sampling seed, no CI; only the temperature cliff (168/300) is robust. |
| 027 | ok | 29,029 SONAR sentences in a 4-domain x 3-length grid, 2500 per cell (web-short 1529); per-point local ID with k in {10,20,50}; synthetic calibration blobs dim {5,10,25,50} at N=7500 | none (frozen SONAR encoder; no model trained) | 1 encode (seed 20270802); synthetic noise band from 5 independent blobs per dim | synthetic 5-seed noise band (2*sd ~0.11) vs stratum spreads 25-38; Moran's I permutation p=0.005; TwoNN bootstrap CI [32.3,33.6]; estimator concordance Spearman 0.8 (participation ratio); density confound r=0.71 reported | GPU0 encode + kNN 63 s, CPU estimators ~19 s; GPU type not stated | SONAR (frozen) | Large n (29k) with calibration, permutation and bootstrap; the claim is deliberately restricted to ordering because MLE magnitudes are off-scale and density-confounded (r=0.71). |
| 027 | ok | 2000 seed sentences iterated up to 30 steps (all settled by step 6); 200-seed stratified subsample judged for semantic drift; 3 fixed points hand-verified | none (frozen SONAR encode/decode) | 1 (deterministic greedy map; seed 20210802 for corpus) | none stated (counts are exact for a deterministic map: 1997/2000 fixed points, 2000/2000 distinct); judge 0 parse failures, 6/6 smoke controls; no CI on the 95% faithful rate | GPU0 full run 84 s; gpt2 perplexity on box; local codex judge; GPU type not stated | SONAR (frozen) | Headline is an exact count over 2000 deterministic trajectories; only the 95%-faithful drift figure (200 judged, no CI) is a sample estimate. |
| 028 | ok | 1500 sentences (021 corpus, 500 per length bin) decomposed with one SAE; 16,384 atoms inventoried; top-10 atoms spot-checked | SAE: in-house BatchTopK h=16384 k=32 trained on cpool SONAR embeddings (w40 ckpt, epoch 40, seed 0; training-set size not stated in RESULT); nothing trained in this row | 1 SAE checkpoint (seed0); deterministic per-sample top-k | none stated (Pearson r only; p<1e-6 quoted for the r=-0.127 fabrication link); no bootstrap CI | GPU0, 1.7 s compute; GPU type not stated | SONAR (frozen) + in-house SAE (trained, one seed) | r=-0.78 on n=1500 has negligible sampling error even without a CI; caveats are one SAE seed/width and length-density collinearity, not n. |
| 029 | thin | 1500 (orig, greedy-recon) pairs with 54 fabrication positives; gate-blind subset 1460 with 35 positives; ~21 test positives per CV fold; reconciliation corpora rp_real autodec/tae (sizes not stated) | logistic stacks on 60% stratified splits; NLI models frozen (nli-deberta-v3-large, bart-large-mnli) | 3 CV split seeds; 1 SONAR checkpoint; 2 NLI models | AUC / P@R0.8 as mean over 3 seeds (P@R0.8 sd ~0.10 stated); no bootstrap CI on the 0.478 AUC | GPU0, NLI over 1500 pairs ~1-2 min; GPU type not stated | SONAR (frozen) + frozen NLI cross-encoders | Below-chance AUC rests on 54 positives (35 gate-blind) with no CI; RESULT flags 'tiny positive count' and high CV variance; reconciliation clean-side is corpus-dependent (0.47% vs 3.85%). |
| 030 | ok | 30,742 decoded tokens from 1500 greedy reconstructions (content 17,146 / function 10,034 / punct 3,562); 384 fabrication tokens (1.25%) for the abstention analysis | PAV/isotonic recalibration fit in-sample on the 30,742 tokens (89 blocks); no held-out split | 1 (deterministic re-analysis of 022 scores) | none stated (no CI on ECE 0.070, AUC 0.738, or Brier deltas); tokens are clustered within 1500 sentences but not cluster-resampled | CPU only, local; no GPU | SONAR (frozen) | Very large token n; direction of miscalibration is consistent across every bin; caveats are in-sample PAV fit and only 384 fabrication tokens for the abstention curve, not the headline. |
| 032 | very thin | 200 pairs x 9 points x 2 methods = 3600 greedy decodes for fluency/geometry; judged coherence on 48 interior points per method (96 total); betweenness vote on 60 matched pairs; hand check 15 | none (frozen SONAR; kNN graph k=15 over 29,029 cached z) | 1 (pair seed 32, judge subsample seed 320320) | none stated (no CI on 0.625 vs 0.500 coherent or 0.042 vs 0.167 word-salad; no CI on mean_lp curves); judge 0/156 parse failures; hand 13/15 | box GPU decode of 3600 points + re-encode; local codex judge; GPU type/time not stated | SONAR (frozen) | The headline percentages rest on 48 judged points per arm (word-salad 2/48 vs 8/48) and a 60-vote preference, single judge, no CI; the 3600-point fluency/nn_decode curves are the better-powered but non-headline evidence. |
| 033 | thin | 40 held-out test pairs per transform (7 base types + compose), disjoint vocabulary; 1240 judged decodes (baseline/predicted/ceiling/symmetry); voice argument-swap 0/40; hand check 15 | offset = mean difference over 60 templated train pairs per transform (frozen SONAR; no model trained) | 1 stimulus seed (33); 1 SONAR checkpoint; single judge pass | none stated (no CI on success rates); gates G_ceiling 0.998 / G_baseline 0.00; diff_align reported as a z-side diagnostic | phys GPU2 decode 30.5 s (1240 jobs); local codex judging; GPU type not stated | SONAR (frozen) | Only 40 test pairs per transform and no CI, but the effects sit at the ceilings (1.00 vs 0.00) with baseline/ceiling gates, so the ordering is robust; templated stimuli, single judge. |
| 034 | thin | 56 held-out natural affirmatives for ADD (peak 0.84 at alpha=1.5; subcategories declarative 39 / short 13 / long 3 / question 1); 40 naturally-negative for UNDO; 688 decode jobs, 744 judged; hand check 15 | offset_nat = mean difference over 100 codex-generated natural train pairs; offset_tmpl refit from 033's 60 templated pairs (frozen SONAR; no model trained) | 1 stimulus seed (34); 1 random-direction null; 1 SONAR checkpoint | none stated (no CI on flip rates); controls: norm-matched random direction 0.00, baseline 0.00, ceiling 0.93; hand 14/15 | phys GPU0 decode 90 s; local codex judging; GPU type not stated | SONAR (frozen) | 56 test sentences and no CI (0.84 = ~47/56), long/question subcategories n=3/1, single judge; the random-direction control (0.00) makes the causal direction claim credible, and the 'best-calibrated' label refers to Brier, not n. |
| 035 | ok | templated pairs: 120 train per region per transform (4 length regions x 5 transforms); 120 held-out test pairs per transform per arm; 2400 decode jobs, 1200 judged; hand check 8/8 | regional and global offsets = mean differences over 120 (regional) / 480 (global) templated train pairs; frozen SONAR | 1 stimulus seed (35); 1 SONAR checkpoint; k-means k=6 as partition robustness | 1000x bootstrap of within-region offset direction (noise band ~3 deg) vs cross-region angle 20-25 deg; no CI on decode success deltas; local-ID correlation uses only 4 region points | phys GPU0 decode 98 s (2400 jobs); CPU geometry; local codex judge; GPU type not stated | SONAR (frozen) | Curvature claim is bootstrap-banded (6-8x ratio) and reproduced under a second partition; the null on application (delta +0.8 pts) is at ceiling on 120 pairs/transform; caveat: templated short sentences only, 4 regions for the ID correlation. |
| 036 | thin | Corpus B 1,500 sentences (headline partial r 0.41 rests on this); Corpus A 29,029 sentences (raw correlations); causal scaling 150 sentences x 5 scale factors = 750 greedy decodes; codex judge 30 items x 3 factors = 90 | none (frozen SONAR; correlations and partial correlations only) | 1 | none — no CI/bootstrap on any correlation or on the scaling null; length-control partial correlations only; judge counts (30/30 'same') reported raw | GPU0 (box A4000) for 750 greedy decodes; CPU for correlations; wall not stated | SONAR (frozen) | Correlation n is fine (1,500 / 29,029) but nothing has a CI; the 'scaling changes nothing' null rests on 150 sentences and a 30-item-per-factor judge with no uncertainty. |
| 037 | thin | Corpus A 29,029 sentences (isotropy statistics — the headline); binding survival: 350 propositions / 3,500 items (n_train 700 / n_test 700 per cell); Corpus B 1,500 (norm-specificity); local-ID 6,000-point stratified subsample | linear role probe on 700 items per cell (battery subsample); frozen SONAR | 1 | binding.json records n_boot 400 but RESULT reports point AUCs only (no CIs); no CI on anisotropy statistics or partial correlations | CPU-only, ~45 s wall (no GPU) | SONAR (frozen) | Isotropy numbers rest on 29,029 points (fine), but the four 'survives' verdicts are single-seed, 350-proposition subsample, no CIs reported. |
| 038 | ok | 400 seed sentences (antipode judged in full n=400; geometry n=400 per condition); random and mean-reflection controls judged on 200 each; z-gate 20; hand-check 15 | none | 1 | chi-square(3) on judge class distributions (antipode vs random 1.16, NS; vs meanrefl 4.57, NS); no bootstrap CI on geometry means; no positive control for the negation_opposite label | GPU0 (box A4000), 1,600 greedy decodes; wall not stated | SONAR (frozen) | 400 items with a chi-square test against a 200-item random null; the geometric readouts (selfconsist 0.097 vs 0.101) are far from the z baseline (0.993) so power is not the issue. |
| 039 | thin | 1,000-point random subsample of 29,029 z (global topology); 6 attribute families of 7–12 values x 24 templates = 168–288 sentences each (1,464 total); planted-circle sanity n=300 | none | 1 | covariance-matched Gaussian null p95 band from 10–15 bootstrap draws; label-shuffle null for closure ratio; split-half subsample stability (two disjoint halves) | CPU TDA ~30 s; GPU0 encode <2 s | SONAR (frozen) | A null claim whose null bands come from only 10–15 resamples and whose attribute clouds are 168–288 points each; the planted-circle positive control (9.3x dominance) is the main defence. |
| 040 | thin | 1,500 orig/reconstruction pairs with 54 labeled fabrications (cosine-gate AUC headline); SAE axis n=1,500; geodesic 150 paths; transforms fit on 29,029 z | none (whitening/PCA fit on 29,029 z; w40 SAE reused, not retrained) | 1 | none — no CI on AUCs; RESULT concedes the ±0.01 AUC deltas on 54 positives are 'well within resampling noise'; 5 random rotations (rot0..4) as control | CPU-only, 115.6 s | SONAR (frozen) + in-house w40 SAE (reused) | The headline null (no basis fixes the gate) rests on 54 positive fabrications with no CI; a +0.05 AUC gain cannot be excluded with that many positives. |
| 041 | very thin | 4,000 val sentences (FVU; 13,706 firing atoms classified into shared/local); 36,000 train sentences | 36,000 pile-10k sentences; crosscoder h=16384 k=32, 250 epochs; plus 3 single-rep TopK SAE baselines | 1 | none (no CI, no seed repeat) | GPU0 (box A4000), ~1,800 s wall | crosscoder / SAE trained in-house on SONAR (frozen) layer states | Single training seed, no repeat, no CI; row 044 later showed same-width cross-seed atom match is only ~15%, so the 26/13,706 shared count is a one-seed number of a seed-sensitive quantity (L24 verdict rests on 33 atoms). |
| 042 | ok | 1,000 val sentences greedy-decoded x 5 conditions (chrF +25.3 headline); 4,000 val for FVU/ICA/AE; 36,000 train | 36,000 sentences (fresh TopK SAE h=8192 k=32 40 ep; shallow AE; FastICA on 20k subsample); w40 SAE reused | 1 | paired bootstrap 95% CI on chrF deltas over the 1,000 decoded sentences (resample count not stated); ICA vs 3-draw Gaussian null band; SAE cell 2 null seeds; single seed for most fits | GPU0 (box A4000), ~8 min wall (encode 116 s, analyze 212 s, decode 119 s) | SONAR (frozen) + SAEs trained in-house | Headline decode lift is measured on 1,000 items with a paired-bootstrap CI [24.10, 26.43]; the '69x' ICA figure rests on a 3-draw null band and the fresh-SAE cell is INSTRUMENT_FAILURE. |
| 043 | thin | Corpus A 12,000 sentences, 70/30 split -> 3,600 test (toklen/BoW/domain probes); Corpus B 12,000 battery items (family/voice/role); BoW pool 66 words | ridge/logistic probes on 8,400 (70%) of 12,000; w40 SAE reused (not retrained) | 1 | none — RESULT: 'Single seed, single SAE; no CI/bootstrap on probe metrics'; shuffled-label baselines only | CPU only, full run 17 s | SONAR (frozen) + in-house w40 SAE (reused) | Test n (3,600) is adequate and the gaps are large, but there is no uncertainty of any kind and a single seed/SAE; the role cell is INSTRUMENT_FAILURE. |
| 044 | thin | 36,000 pile z for atom firing/matching (2,036–15,597 narrow atoms per width pair); probes on corpus-A 12,000 (3,600 test); FVU on cpool-val 20k / pile-val 4k / corpus-A 12k | 2 new h2048 SAEs on 280,000 (w40 c_pool), 40 epochs; 8 reused w40 ckpts (h8192–65536 x 2 seeds) | 2 (s0,s1 per width for cross-seed stability); matching and probes on seed 0 only | bootstrap CI (20 half-corpus resamples) on the h16384 cross-seed match fraction only [0.150, 0.151]; no CI on split fractions, FVU or probe metrics | 2x A4000; h2048 trains 48 s each; whole chain ~2 min | SAEs trained in-house on SONAR | Splitting/absorption cells are formally INSTRUMENT_FAILURE (cross-seed match 15% < 20% gate); the surviving FVU/probe headline is seed-0 only with no CI. |
| 045 | ok | 1,000 pile-val sentences greedy-decoded x 12 conditions (chrF headline); 4,000 pile val + 12,000 corpus-A rows for FVU / atom-mass | none (w40 SAE h16384 k32 seed0 reused; inference-time m only) | 1 (3 random draws for the rand-m control) | paired bootstrap 95% CI on chrF deltas vs M32 (n=1,000; e.g. M128 +14.41 [13.45, 15.37]); no CI on FVU | GPU0 (box A4000), 4 min (decode 235 s) | SONAR (frozen) + in-house SAE (reused) | 1,000 decoded items with paired-bootstrap CIs on every ladder step; FVU on 4k/12k rows; single SAE and m>32 is off-distribution (flagged). |
| 046 | ok | 2,009 FLORES-200 parallel rows x 6 languages (dev 997 + devtest 1,012); matcher candidate atoms 268–302 per language pair | none (w40 SAE reused) | 1 SAE seed; 20 derangements for the null | scrambled-alignment null = max of 20 derangement medians (Jaccard); matcher certified by self-match / noise / derangement controls; identity rate saturates at 1.000; no CI | GPU0 (box A4000), 58 s (encode 49.5 s) | SONAR (frozen) + in-house SAE (reused) | 2,009 parallel rows, controls at full n, and a saturated identity rate with positive minimum margin in every pair; single SAE seed is the main caveat. |
| 047 | thin | 5,000 FrameNet sentences (50 frames x 100), 70/30 stratified split -> 1,500 test; 46/50 frames selective; 1,961 atoms firing >=20 | 3,500 sentences for centroid/PCA bases and 50-way logistic probes; w40 SAE reused | 1 | 20 label-permutation null for atom-frame F1 / MI; no CI on FVU (P2 margin 0.017 'un-bootstrapped') or on probe accuracies | GPU0 encode 14.6 s; ~3 min chain | SONAR (frozen) + in-house SAE (reused) | The FVU ordering that carries the headline (frame 0.839 vs SAE 0.822 vs PCA-50 0.770) has no CI and a 0.017 margin on one seed; 1,500 test sentences. |
| 048 | thin | 20,000 owt_val eval rows for activation-correlation matching; 512 atoms per SAE; 5 checkpoints per organism | TAE ladder rung-A A_P90 (d_z=256, 12M params, 500M-token budget), 2 organism seeds, checkpoints pre-existing; 20 TopK SAEs (h=512 k=16) each on 320k owt_train lines, 120 epochs | 2 organism seeds x 2 SAE seeds (cross-SAE-seed primary) | none (no CI); organism-s1 replication within 0.03; ceiling-normalised by cross-seed match; scrambled null 0.000 | 2x A4000 in parallel, 8.4 min | TAE ladder (trained in-house) + SAEs trained in-house | Replicated across 2 organism seeds and ceiling-normalised, but no CI, only 5 time points with a 4000->26000 gap, and the SAE config was recalibrated pre-run after the frozen one proved non-identifiable (ceiling 0.008). |
| 049 | ok | 5,000 pairs per cell x 6 cells = 30,000 pairs (PAWS + QQP); overlap-matched deciles >=100 pairs/side; per-atom analysis on 414 candidate atoms | none (w40 SAE reused) | 1; 20 derangements | scrambled-pairing null (max of 20 derangement medians); overlap-matched decile comparison (10/10); no CI on medians | GPU0 (box A4000), 165 s (encode 135 s) | SONAR (frozen) + in-house SAE (reused) | 5,000 pairs per cell with a derangement null and matched-overlap deciles; no CI but the item count makes the medians stable. |
| 050 | ok | 8 w40 training runs x 40 epochs audited (320 epochs); 2,232 z-dead / 46,106 z-rare / 17,198 z-healthy atoms (h65536) on pile 40k; cpool chunk0 100,054 rows; corpus A 12k; 40 atoms/class text necropsy | Arm C: 3x h16384 2-epoch mini-retrains on 280k (aux256 x2, aux0); original 8 runs 280k x 40 ep (reused) | 2 (s0,s1) for the audit; Arm C 1 seed + determinism envelope | class-permutation tests on geometry (p=5e-4); size-matched random-atom baseline; Gaussian-direction null; no CI | CPU (Arms A/B) + GPU0 (Arm C); 3 min 53 s | in-house w40 SAEs (on cpool / SONAR z) | The 'aux never engages' claim is an exhaustive audit of all 8 runs plus a bit-identical ablation; the necropsy compares 2,232 atoms with permutation tests. |
| 051 | ok | agent_patient 2,000 items (200 propositions x 2 orders x 5 families) + genitive 1,200, per model x 6 models; cross+lex-holdout cells have n_test 100–400 each, pooled | linear + 1-hidden MLP probes on battery train split (n_train 300–400 per cell); 6 frozen external embedders | 3 readout seeds (0,1,2) | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU2 (box A4000); per-model wall not stated | external models (MiniLM, mpnet, e5-base, bge-base, gte-base, sentence-t5-base) | 2,000 items, 3 seeds, clustered bootstrap; the row honestly reports INSTRUMENT_FAILURE for all six (no null certified). |
| 052 | thin | battery 2,000 + 1,200 items per model (2 models); secondary 061 case-marking battery 1,500 stimuli/lang (deu, eng) per model | probes on battery train split; external LASER2 (44.6M) and LaBSE (471M), frozen | 3 readout seeds | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys0 (A4000), ~7 min | external models | Item n and CIs are fine, but the causal headline ('it's the decoder') rests on one model per objective family (n=1 vs n=1) with pooling confounded, and the LaBSE 0.986 possessor cell is post-hoc. |
| 053 | ok | battery 2,000 + 1,200 items x 6 conditions x 2 models = 12 batteries | probes on battery train split; external instructor-base and multilingual-e5-large-instruct (560M), frozen | 3 readout seeds | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys0 (A4000), ~1.6–2.1 min per battery | external models | 2,000 items with clustered CIs in every cell; caveats are two models and one instruction phrasing per role, and P3 landed 7e-5 from its threshold. |
| 054 | ok | battery 3,200 sentences (2,000 + 1,200) x 4 representations; role-swap cell 100 contexts x 2 variants = 200 generations, all 200 judged | probes on battery train split; Mimir-1.6B (external LCM reproduction; Meta weights unavailable) | 3 readout seeds; 1 diffusion sample per context (seed 42) | proposition-cluster bootstrap 95% CI, n_boot 1000 (battery); exact binomial on forced choice (n=200, p=0.856); judge counts without CI | GPU phys0 (A4000), ~9 min | external model (Mimir-1.6B two-tower diffusion LCM) over SONAR space | Battery headline has 3,200 items and clustered CIs; the 200-generation role-swap cell's preregistered metric was invalidated and the judge cell is an upper bound. |
| 055 | ok | battery 3,200 sentences x 8 GSM8k representations (emb_mean, base_h, thought1–6); functional gate 60 GSM8k questions; ProsQA arm (7 batteries) never ran | probes on battery train split; external community GPT-2 Coconut checkpoints (gsm8k checkpoint_33, prosqa checkpoint_40) | 3 readout seeds; 1 checkpoint per arm | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys0 (A4000), ~14 min | external model (GPT-2 Coconut reproductions) | 3,200 items with clustered CIs on the GSM arm; half the design (ProsQA) was refused by its manipulation gate so P5 is unscoreable, and only one checkpoint per arm. |
| 056 | ok | R-decl 8,000 forced choices (4,800 surface-incongruent = primary); R-qa 2,000; L-lex 2,000; S-scram 2,000; localization on 2,000 R-decl + 1,000 L-lex subsamples; head cell 600+600; CLS battery 3,200 | none for behavioral cells (zero-shot deterministic scoring); battery probes on train split | 1 (deterministic scoring, seed-0 subsamples); battery 3 readout seeds | none on behavioral accuracies (no CI/bootstrap stated); battery cell has cluster bootstrap n_boot 1000 | GPU phys0 (A4000), ~2 min + battery 86 s (~300k cross-encoder forwards) | external models (cross-encoder/ms-marco-MiniLM-L-6-v2 vs all-MiniLM-L6-v2) | 4,800 items behind the below-chance primary; no CI stated but n is large and lexical control is 1.000; single reranker is the generalization caveat. |
| 057 | ok | battery 3,200 sentences x 8 representations (6 masking rates + 2 anchors) = 32 rep x readout cells; functional gate 20 ~100-token paragraphs | probes on battery train split; external DiffuGPT-small (124M), frozen | 3 readout seeds; K=4 deterministic masking draws averaged | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys0 (A4000), ~13 min | external model (DiffuGPT-small masked-diffusion LM) | 3,200 items, clustered CIs in all 32 cells; single small model and final-layer-only are scope caveats, not sample-size ones. |
| 058 | thin | 3,200 unique battery sentences (alignment cos / retrieval / battery); cross-modal decode 200 sentences x 3 representations; G0 certification 8 sentences + 17 words | none (frozen SONAR speech and text encoders; probes on battery train split) | 3 readout seeds; 1 deterministic TTS voice | proposition-cluster bootstrap 95% CI, n_boot 1000 (battery); none on alignment cosine, retrieval P@1, or decode chrF/EM (n=200) | GPU phys0 (A4000), ~15.5 min (TTS 185 s CPU, speech encode 80 s) | SONAR (frozen; speech + text encoders, text decoder) | Battery part is well-powered, but the decode headline (chrF 89.7 vs 92.9, 'zero fabrication') rests on 200 sentences, no CI, one synthetic TTS voice. |
| 059 | ok | battery 2,000 + 1,200 items x 4 scales (gtr-t5-base 110M, large 335M, xl 1.24B, xxl 4.8B) | probes on battery train split; 4 external GTR-T5 models, frozen | 3 readout seeds | proposition-cluster bootstrap 95% CI, n_boot 1000 | one A4000 16GB; battery walls 97/115/144/259 s (~10.2 min total), xl/xxl in fp16 | external models | 2,000 items, 3 seeds, clustered CIs at every scale; the internal genitive all-vocab control shows the instrument can see a trend. |
| 060 | very thin | 100 rating pairs (5 cells x 20 pairs) + 30 forced-choice triplets; 3 LLM panelists x 3 conditions = 900 ratings + 90 triplet judgments (1,230 judgments total); within-topic correlation on 80 pairs | none | 1 (seed-60 pair build; 3 deterministic CLI panelists, 1 wording-variant each) | bootstrap 95% CI on the Spearman difference only (P5: [−0.097, +0.002]); no CI on cell means, the 4.04 gap, or the −0.218 within-topic correlation | local CPU only, 51 CLI calls, ~13 min; SONAR z reused from row 058 cache | external LLMs (codex, grok, Gemini-backed) as proxy raters + SONAR (frozen) z-cosine | 20 pairs per cell and 30 triplets, one template register, LLM proxies rather than humans (row's own T3-LLM-PROXY ceiling); the 90/90 vs 30/30 split is unanimous so the inversion is not a power artifact, but its basis is tiny. |
| 061 | thin | 1,500 templated stimuli per language x 4 languages (750 propositions x 2 word orders); primary cross-order + lexical-holdout cell n_test 250 per language (within-role cell n_test 750) | linear probe, n_train 500 (cross-lex) / 750 (within); frozen SONAR | 3 readout seeds (per-seed primary AUC identical — deterministic linear fit on frozen split) | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys3 (A4000), 51 s wall | SONAR (frozen) | German's positive is CI-lo 0.608 on a 250-item lexical-holdout test cell, linear-only, one construction pair per language; row 062 then failed to replicate it on fresh vocabulary (0.551). |
| 062 | thin | 1,500 sentences per language x 4 languages (6,000 total): 500 train + 250 test propositions x 2 orders; test cell 500 items/lang; 061-replication cell on the 250 held-out propositions | linear probe on 1,000 train items per language; frozen SONAR | 3 readout seeds | proposition-cluster bootstrap 95% CI, n_boot 1000 | GPU phys0 (A4000), 30 s wall | SONAR (frozen) | Primary transfer cells are INSTRUMENT_FAILURE by design flaw; the surviving headline (jpn 0.696, deu 0.551, jpn->X ~0.45) rests on one 250-proposition fresh lexicon, linear-only, single construction pair. |
| 063 | thin | 64 items per (language, c) cell x 6 concatenation lengths x 6 languages = 2,304 items (384 per language); each language's knee estimated from its 6 bins of 64 | none | 1 (greedy deterministic; s0–s2 files are GPU shards, not seeds) | ci95 per (language, bin) on mean chrF in results JSON (method/resamples not stated in RESULT); no CI on knee location; the offset-vs-knee Spearman −0.6 is on n=5 languages (post-hoc) | 3x A4000, 9.3 min wall (decode ~7 min) | SONAR (frozen) | Each per-language knee is a frozen-criterion crossing between 6 bins of 64 items with no CI on the knee itself; chrF cross-script confound flagged by the row. |
| 064 | ok | offset constancy on all 1,012 FLORES devtest rows x 5 languages; steering on the first 200 devtest sentences x 5 languages x alpha in {0,0.5,1,1.5,2} x 2 decoder tokens | v_L = mu_L − mu_eng estimated from FLORES dev (997 rows) — a mean, no model trained | 1 | 95% CI on chrF delta (alpha=1 vs 0) per language (e.g. deu +1.58 [−0.05, 3.23]; bootstrap, resamples not stated); flip rate 0.000 is threshold-free over 200 x 4 alpha x 5 langs; no CI on EV_const | 3x A4000, 5.7 min wall | SONAR (frozen) | 1,012 rows for the constancy result and 200 sentences per language with CIs on the steering deltas; 0% flip over 4,000 steered decodes is not a power issue. |
| 065 | thin | 300 code-switched items per (language x pattern) cell, 9 cells = 2,700 items, each decoded with 2 tokens; interpolation 200 rows x 3 languages x 5 t x 2 tokens; codex judge 12 items/cell (108) | none (Dice alignment lexicon over 2,009 FLORES parallel pairs) | 1 | none stated — no CI on placement medians, round-trip chrF, or insertion-fate rates; judge on 12 items/cell; ~5% detector false-mixed floor | 3x A4000, 5.8 min wall (~12.9k greedy decodes) | SONAR (frozen) | 300 items per cell is reasonable but no uncertainty is quantified anywhere, the judge sees 12 items/cell, and alignment is a noisy heuristic. |
| 066 | thin | 64 items per (language, c) x 6 lengths = 384 items per language x 3 languages (eng, deu, jpn) = 1,152; per-language bits curve and ceiling from 6 bins of 64 | none (teacher-forced decoder likelihoods) | 1 | none — no CI/bootstrap (results JSON has no se/ci fields); shuffled-z null certifies the instrument; the ~460-bit specific ceiling is post-hoc with a roll-1 null the row says likely overstates the generic component | single A4000, 3.7 min forward-only | SONAR (frozen) | Bits per bin are means over 64 items with no CI, and the headline ceiling (~460 bits) is a post-hoc correction not in the prereg (P2/P3 scored FALSE). |
| 067 | very thin | 192 owt_val chunks x 7 nested prefix lengths (8–62 tokens) per model; knee per model from 7 bins | 3 new TAE trains (ladder rung-A arm-D, d_z 16/64/1024, full 500M-token budget, ~37,360 steps each) + 2 reused d_z=256 runs (s0, s1) | 1 per new dimension; 2 at d_z=256 only | none (no CI); the d=256 seed gap (knee 0.37 tok, I_spec 5.2 bits) serves as the only noise anchor | 3x A4000 in parallel, ~75 min wall (~1.7 GPU-h per train, ~5064 s each) | TAE ladder (trained in-house) | One training seed per new bottleneck size, 192 eval items, no CI; the frozen knee criterion is degenerate at d_z=16 so the scored verdict (objective-limited) and the post-hoc reading (capacity-limited) disagree. |
| 068 | thin | 64 items per (tier, bin) x 7 token bins x 4 tiers = 1,792 items; each tier's knee from its 7 bins of 64; fabrication check ~36 pairs (qualitative) | none | 1 | ci95 per (tier, bin) on mean chrF in results JSON (method not stated in RESULT); no CI on knee_tok or bits-at-knee; frozen gates G1 (length matching) and G2 (perplexity separation) both FAILED | 3x A4000, 7.4 min wall; gpt2 for perplexity | SONAR (frozen) + external gpt2 as perplexity reference | The '1.08x constant bits' headline is three point-estimate knees (one per natural tier) from 64-item bins with no CI on the knee; two frozen gates failed. |
| 069 | thin | pairs per tier: UNREL 65 / REL 80 / REDUN 96 / SELF 67 / LONG 85 (393 pairs; each yields z_A, z_B, z_AB, z_BA) | none (teacher-forced likelihoods; REDUN paraphrases produced by SONAR round-trip) | 1 | 1000x pair bootstrap 95% CI on R (ratio of means) per tier (e.g. UNREL 0.924 [0.911, 0.937]); no CI on cross-carry bits or per-clause fidelity | 3x A4000, 3.8 min wall | SONAR (frozen) | 65–96 pairs per tier (under ~200) though with a 1000-resample pair bootstrap and tight CIs; single seed, English only, one join template. |
| 070 | very thin | 48 items per cell (digit-count x number-type x context; e.g. INT 9 digit levels x contexts, DEC/WORD/YEAR/PHONE cells) from 8 templates per type, + 120 FLORES real-number sentences; every headline exact-match figure is one 48-item cell | none | 1 (8 templates per type) | exact_ci and I_spec_NUM_ci per cell in results JSON (method/resamples not stated in RESULT); codex judge skipped; failure taxonomy counts are raw (e.g. INT_d12_K: 15 exact / 15 rounded / 17 magnitude) | 3x A4000, ~3 min wall | SONAR (frozen) | Each headline number (INT-S D=12 exact >=0.96; INT-K D=12 0.31) is a single 48-item cell built from 8 templates, single seed; the monotone trend across cells is coherent but no cell exceeds n=48. |
| 071 | thin | 48 items/cell x 15 cells (N in {1,2,3,4,5,6,8} x MATCHED and NATURAL regimes = 672 templated items) + REAL cell of 120 FLORES sentences; the '~3 entities at fixed length' headline rests on the 7 MATCHED cells of 48 items each; codex judge on 40 rule-ambiguous items (non-gating) | none (frozen SONAR); offsets/bits are teacher-forced, no probe trained | 1 | none: per-cell point rates and mean I_spec bits, no CI, no bootstrap, no permutation test | 3 GPUs on the 4l box (type not stated in row), ~3.7 min full run after smoke; frozen SONAR decode + teacher-forced bits | SONAR (frozen) | Each cell is 48 items with no uncertainty quantified; the plateau-at-~3 reading compares 48-item cells whose rates differ by <0.1, and G1 passed only at 0.875. |
| 072 | thin | 48 items/cell x 6 demand levels K=1..6 (288 items; each item carries K six-role facts, so per-role spans at K>=4 pool 3 cells x 48 items x K facts); headline deletion order = mean PRESENT survival over K>=4 | none (frozen SONAR); numpy OLS survival~role+rarity+logtok as a control, not a trained model | 1 | none on survival rates (no CI/bootstrap); Spearman rho over 6 roles; OLS partial-out for rarity/length | 3 GPUs on the 4l box (type not stated), ~6.5 min after smoke; frozen SONAR | SONAR (frozen) | 48 items/cell and no CIs; the PATIENT-first gap is large (0.055 vs 0.25-0.38) but the AGENT~PLACE~ACTION ordering is within a few points with no error bars, and the G1 positive-control gate FAILED as coded (0.688 vs 0.85). All 6 predictions false. |
| 073 | thin | 2000 stimuli_v2 agent_patient items per checkpoint; 18 checkpoints = 2 arms (A_D plain DAE, A_P90 paraphrase) x 2 seeds x 4-5 milestones (A_D grid: steps 0/2000/4000/final only) | existing ladder rung-A checkpoints (~12M params, d_z=256, 500M-token budget); no training in this row; battery probes are linear/MLP with 5-fold CV | 2 training seeds per arm (4 trajectories); probe seeds 0,1,2 | battery cluster bootstrap n_boot=1000 (proposition-clustered) on each checkpoint AUC; no CI on onset/saturation step; G1 reproduces baseline to 4 dp | CPU-only (CUDA_VISIBLE_DEVICES=''), 18 checkpoint batteries in ~15 min | TAE ladder (trained in-house, rung A) | Per-checkpoint AUCs are well powered (2000 items, bootstrap) but the trajectory has only 4 milestones on the primary A_D arm with an 11%->100% gap; the 'gradual' claim rests on the other arm's 5-point grid; 2 seeds. |
| 074 | very thin | 2000 stimuli_v2 agent_patient items x 15 checkpoints (battery); val on 20k held-out OWT (val_f1/exact); train positive control on a fixed 512-sample of the memorized set | 6,000 OWT sentences (memorized set), 2.5e9 target tokens = 15,151 epoch-equivalents, rung-A arm-D d_z=256 ~11.6M params, batch 512, lr 2e-4 cosine (v2 amendment) | 1 | battery bootstrap 95% CI (n_boot=1000, proposition-clustered) per checkpoint; frozen grokking criterion uses non-overlapping CIs; no cross-seed | 1 GPU (phys0, type not stated; 075 prereg identifies box GPUs as A4000), ~5.6 GPU-h train (v2 run 10:29-16:05Z) + ~14 min CPU battery; a v1 run was killed at 4.5% budget by a watchdog and discarded | TAE ladder (trained in-house, rung A) | A null (no delayed emergence) from a single seed, single rung, single over-training run with no repeat; the design also holds lexical diversity low (6k sentences), which the RESULT admits cannot separate 'needs diversity' from 'needs diversity and time'. |
| 075 | ok | 8 seeds (6 newly trained s2-s7 + reused A_D_s0/s1); per seed: knee on 192 owt_val chunks x 7 prefix lengths, 5 offset operators on 033 stimuli (Procrustes anchor set n=582 per 077), battery 2000 items | 8 full-budget rung-A arm-D d_z=256 runs, 5e8 tokens each, batch 512, lr 5e-4, matched args (step 37359+-1) | 8 | cross-seed CV (std/mean) over 8 seeds; battery bootstrap n_boot=1000; operator cosines are all-vs-reference-seed medians (not full pairwise, no CI) | 6 new runs x ~1.08 GPU-h on RTX A4000 (~6.5 GPU-h) + ~0.4 GPU-h eval; 2 waves of 3 GPUs, ~2.6 h wall | TAE ladder (trained in-house, rung A) | 8 independent training seeds is the strongest replication in this range; caveat is a single rung/arm/d_z cell and operator geometry compared only against one reference seed. |
| 076 | thin | 5 curricula (random_a, random_b, short2long, long2short, common2rare), one run each; per run: knee on 192 owt_val chunks, battery 2000 items, 5 operators; G3 multiset hash identical across all 5 (19,900,000 lines each) | 5 runs x one finite pass over the same 5e8-token pool (19.9M lines), rung-A arm-D d_z=256, seed 0, lr 5e-4 cosine, batch 512; long2short re-run once after watchdog removal | 1 | cross-order CV compared against 075's 8-seed CV band; random_a-vs-random_b reshuffle serves as the order-null floor; battery bootstrap n_boot=1000; no repeat of any curriculum | 5 x ~0.9-1.1 GPU-h (RTX A4000) ~5-5.4 GPU-h + ~15 min CPU pool build; 3 GPUs, ~2-2.5 h wall (+ long2short re-run) | TAE ladder (trained in-house, rung A) | Each curriculum is a single seed-0 run with no repeat; the reconstruction collapse (val_f1 0.11-0.52 vs 0.645, ~104x the seed band) is far too large to be seed noise, but the surviving-arm operator/surfX comparison rests on only the 2 reshuffles plus a partially collapsed arm; mechanical aggregator gave the opposite verdict (0.173) until re-scored by hand. |
| 077 | very thin | 4 conditions (DAE-only, MT-only, DAE->MT, MT->DAE), one run each; per run: knee 192 chunks, battery 2000 items, 5 operators (anchor n=582) | 4 runs x 5e8 tokens (switch at 250M), rung-A d_z=256, seed 0, single continuous cosine LR; MT arm = ParaNMT monolingual paraphrase, not bitext | 1 | memory-of-first index m per invariant; comparison against 075 seed band (std val_f1 0.003, surfX 0.019) imported, not re-measured; battery bootstrap on cells; no repeat | ~2.8 GPU-h train + ~0.5 GPU-h eval (4 runs x ~45 min, 3 GPUs); aggregator crashed on a np.eye bug and was re-run CPU-only | TAE ladder (trained in-house, rung A) | n=1 run per cell, single seed and single switch point; the surfX hysteresis is 0.72/0.71 vs 0.61/0.60 with an imported seed std of 0.019, all 5 frozen predictions scored FALSE (Brier 0.421) and the mechanical verdict is admitted to be band-brittle; RESULT itself asks for >=2 seeds x 2 switch fractions. |
| 078 | thin | 11 transforms; held-out TEST ~28 items per marker/control transform and 18 per causal transform (700 items total, 1440 decode jobs, 1440 judge records); arg-swap headline pools 0/74; MARKER headline pools 5 x ~28 | offset = mean(z_after - z_before) on lexically disjoint TRAIN pairs (~32 per transform per RESULT's MLP note); small weight-decayed MLP control on the same train set | 1 | none (point success/collateral rates, single local codex judge, no CI/bootstrap); ceiling/baseline gates 0.98/0.00 | CPU decode 576.7 s + local codex judge 684 s; frozen SONAR | SONAR (frozen) | 18-28 held-out items per transform and no CIs or judge agreement; outcomes are saturated (1.00 vs 0.00) and pooled cells reach 74-140 items, so the boundary claim is robust in direction but the per-transform numbers are unquantified; spatial_swap ceiling only 0.80. |
| 079 | very thin | 24 held-out test items per pair x 6 pairs (5 linearizing + 1 control) per decode kind; 384 stimuli, 1152 decode jobs, 1152 judge records | v(A), v(B), v(AB) fit as mean offsets on shared TRAIN base sentences (n not stated in RESULT; disjoint test vocab) | 1 | none (one additivity cosine per pair, point success rates, single codex judge effort=low, no CI) | CPU decode 492.7 s + local codex judge 460 s; frozen SONAR | SONAR (frozen) | Headline additivity (mean 0.982) is 5 numbers, one per pair, with 24 items per decode cell and no uncertainty; RESULT itself calls it 'a small, hand-picked operator zoo'; the control's additivity cosine was degenerate (0.999), so P5 failed on form. |
| 080 | thin | 90 forward + 50 reverse held-out bases per operator x 4 operators (560 items) x alpha grid {-1,-0.5,0,0.5,1,1.5,2} = 5920 decodes; 900-sentence real bank for nn-cos | v_T = mean(z_target - z_base) over 64 TRAIN pairs per operator (TRAIN vocab disjoint from TEST) | 1 | none (rule-based deterministic readout, point success per alpha, no CI; codex fluency cross-check written but NOT run) | GPU phys0 (type not stated), 86.4 s decode; frozen SONAR | SONAR (frozen) | 90 bases per operator with saturated effects (0.00 -> 1.00, random 0.00) is directionally solid, but no CI, rule-based readout only, templated stimuli, and P2 (over-steer) missed. |
| 081 | very thin | n_fwd=40 and n_rev=16 held-out bases per operator per checkpoint; 3 operators x 18 checkpoints (2 arms x 2 seeds; A_D grid 0/2000/4000/final) = 14,688 decode-steering points | existing ladder rung-A checkpoints (no training); offset v fit on 080's TRAIN pairs (64) with each checkpoint's encoder | 2 training seeds per arm (4 trajectories); no probe seeds | none (point success rates per cell, no CI/bootstrap) | CPU-only, 720 s (12 min) | TAE ladder (trained in-house, rung A); rule-based readout from 080 | Every causal-success cell is 40 items (16 for invertibility) with no CI; the install-step claims (2000 vs 4000 vs never) rest on an A_D grid of 4 milestones; consistency across 2 seeds x 2 arms is the only replication; the G_negctrl gate split. |
| 082 | thin | 96 base/flip pairs per class x 8 classes (6 meaning-flip + 2 paraphrase); pooled miss-rate over n=527 validated flips; 2300 encode/decode jobs (meta n_test 720) | none for the primary (real sentence pairs); secondary offset generator fit on TRAIN pairs | 1 | none (per-class mean/median/min/max cosine, no CI; rule-based readout; codex judge NOT run) | GPU phys0 (type not stated), 42.1 s; frozen SONAR | SONAR (frozen) | 96 items per class with no CI; the key ordering (negation min cos 0.908 >= synonym mean 0.906) is stated at the distribution level and is a small margin; templated stimuli; 3/7 predictions correct (Brier 0.334). |
| 083 | very thin | 32 held-out test bases (300-sentence bank, 29 carriers, 12 magnitude fractions = 21,344 sweep decodes); round-trip demo on 384 payloads = 32 sentences x 12 payloads; detectability AUCs computed on the 32-vs-clean split per family | none (PCA on the 300-sentence bank for carrier directions; no probe) | 1 | none (medians per carrier; rank-statistic AUC without CI; capacity is a modelling estimate with gamma=4 margin) | GPU phys0 (type not stated), 227.7 s; frozen SONAR | SONAR (frozen) | The ~103-bit capacity and the 0.92-0.94 / 0.998 detection AUCs rest on 32 test sentences with no CI; capacity depends on the gamma choice and the auc_spanE field is admitted degenerate. |
| 084 | very thin | 48 canary propositions + 48 role-swap twins in a 1596-item store (1500 corpus); 4 query types; 96 round-trip decodes; headline = 48/48 and median twin-gap +0.159 | none (frozen SONAR, cosine retrieval) | 1 | none (counts out of 48, medians; no CI or permutation) | GPU phys0 (type not stated), 12.8 s; frozen SONAR | SONAR (frozen) | Headline rests on 48 templated person-verb-person canaries, 1 seed, no CI; the 48/48 result is saturated (exact binomial lower bound ~0.93) but the +0.159 margin has no spread reported; P2/P4 missed. |
| 085 | very thin | 48 triples x k=0..12 shared-word ladder (13 distractors each) in a 1248-item store; ~2200 unique encodes; headline flip at k=1 = 47/48 (0.979) | none (frozen SONAR; deterministic surface-trusting QA reader) | 1 | none (fractions of 48, no CI) | GPU phys0 (type not stated), encoder-only, 7.5 s | SONAR (frozen) | 48 hand-built worst-case triples, 1 seed, no CI; the verbose 12-topic-word query is itself a poor retriever (unrelated filler beats the gold fact in 65% of triples before any distractor), so the '+0.123 per word' headline is confounded by query dilution, as the RESULT concedes. |
| 086 | very thin | 96 base sentences x 9 categories (7 corruptions + clean re-encode + paraphrase) = 864 z' scored; fusion trained/tested on a 60/40 split of the 96 bases (~58 train / ~38 held-out bases per category); bank 4097 | standardized logistic-regression fusion over 6 scores, trained on ~58 bases x categories (60% split); prototype (nearest-centroid) probes | 1 | none (AUC via rank statistic, no CI; single split; RESULT flags logistic weights as 'small-sample') | GPU phys0 (type not stated), 42.9 s; frozen SONAR | SONAR (frozen) + logistic fusion trained in-house | Held-out AUCs rest on ~38 bases per corruption with no CI; the primary AUC 1.000 is saturated but the deployment-relevant secondary numbers (covert 0.68, negation 0.85, roleswap 0.81) are unquantified small-sample estimates from one split; 3/6 predictions correct. |
| 087 | thin | 80 bases x 5 categories = 400 records (500-sentence bank); 400/400 codex judge labels; headline interp fabrication 92.5% = 74/80, interp reject-AUC 0.508 on 80-vs-80 | none (nn-cos density to a 500-sentence real-z bank) | 1 | none (AUC and fabrication rates as point estimates, no CI; single condition-blind codex judge, no second rater) | GPU phys0 (type not stated), 92.9 s + local codex judge 183 s | SONAR (frozen) | 80 items per category, no CI, single judge; the qualitative split (interp AUC 0.508 vs off-manifold 1.000) is far from marginal, but the coverage/fabrication tradeoff curve and 'cannot go below ~33%' are unquantified; 2/5 predictions correct (Brier 0.330). |
| 088 | thin | 2000 members / 2000 non-members per ladder arm (MEM, STD, UNTRAINED) and 400/400 for the SONAR proxy; 70/30 split so ~1200 held-out per ladder arm, ~240 for SONAR | logistic membership classifier on 70% (2800) per arm; organisms reused: 074 final (memorized 6k set), A_D_s0 (5e8 tokens, seed 0), 074 ck_0 untrained; forward-only | 1 | 10-shuffle label null per arm (degenerate on MEM/SONAR: 0.672/0.590); no CI on AUC; single organism per arm | GPU phys0 (type not stated), 228 s forward-only | TAE ladder (trained in-house, reused organisms) + SONAR proxy | Item counts are ample (4000 per arm) but the STD null (AUC 0.546) rests on one seed-0 organism at ~1.16 epochs with no cross-seed or epoch sweep and no CI; the SONAR arm has no true membership labels. |
| 089 | thin | per attribute (balanced, per_class=2500): register 7500, topic 10000, sentiment 5000, formality 5000, acceptability 4982, age 5000, gender 5000; 70/30 stratified split so ~1500-3000 held-out per attribute | StandardScaler + logistic probe on 70% (~3500-7000 texts) per attribute; TF-IDF BoW baseline on the same split; OLS topic residualization for blog arms | 1 | shuffled-label null per arm (0.483-0.519); no CI on AUC; single split (seed 0) | GPU phys0 (type not stated), 172.6 s forward-only; frozen SONAR encoder | SONAR (frozen) + linear probes trained in-house | Thousands of independent items per attribute (AUC SE roughly 0.01), so the headline ranking is safe, but no CI and one split, which matters for the marginal calls: age 0.717 vs the 0.70 bar (P5 scored FALSE by 0.017) and z-vs-BoW gaps of +0.024/+0.031. |
| 090 | n/a (instrument or theory) | synthetic Anchor A n=1000 d=256; Anchor B n=80 d=768; real anchor 2000 cached SONAR agent_patient z; effect grid d in {0,.5,1,1.5,2,3,4} | kit runs the battery's own group-aware 5-fold CV linear/MLP probe verbatim | 5 | cluster-bootstrap CI n_boot=1000 (battery machinery); MDE spread across 5 seeds = 0.0 | CPU-only, 86 s wall; self-tests ~30 s | instrument (no data claim) | Tooling row validated on synthetic anchors plus one real SONAR anchor; only empirical note is P6: SONAR within-family role AUC is 0.769, not >=0.9. |
| 091 | n/a (instrument or theory) | stimuli_v2 battery, 4 tasks (agent_patient 2000, genitive 1200, causal 2000, temporal 2000 items) plus 6 deliberately broken copies; 11 self-tests | none (pure-stdlib property checker) | 1 | none needed (deterministic PASS/FAIL checks) | CPU, pure stdlib, seconds inline | instrument (no data claim) | Deterministic checker; validation is exhaustive over the shipped battery and injected bugs. |
| 092 | n/a (instrument or theory) | real anchor: 027/036 SONAR corpus n=29029 z (norm-split and norm-orthogonal split) + synthetic magnitude fake, operator-faithful synthetic, control sweep; 13 self-tests | none (numpy linter; 2-fold cross-fit AUC) | 1 | 2-fold cross-fit AUC to avoid dim>>n inflation; no CI | CPU numpy, deterministic; wall not stated | instrument (no data claim) | Tooling row; note the norm-only-AUC gate was added mid-build after the prereg's cosine-kills-it belief was falsified on the real anchor (disclosed). |
| 093 | very thin | 24 codex-generated candidate families (3 rounds x 8); 11 well-formed+plausible families SONAR-tested; each family realized over 48 shared propositions x 2 orders = 96 items; 3 injected artifact controls | linear role probes per family on 96 items (single seed); ppk.certify power gate on pooled baseline | 1 | none (point AUCs per family; RESULT states no cluster-bootstrap CIs on per-family cells) | local codex loop (CPU) + SONAR encode on box GPU0, 23 s | SONAR (frozen) + instrument (adversarial loop) | A null ('zero genuine breaks') over only 24 LLM-proposed candidates (11 tested on SONAR) with 96 items each, single seed, point AUCs without CIs, and the strongest known fragility axis (filler pools/lexical diversity) deliberately excluded from the attack surface; P1 (0.90) missed. |
| 094 | n/a (instrument or theory) | 4 past rows' frozen predictions (069, 085, 088, 039) reproduced; 23 validation checks; 35 self-tests | none (stdlib library) | 1 | none needed (exact arithmetic reproduction to <=1e-4) | CPU, milliseconds | instrument (no data claim) | Bookkeeping library; deterministic. |
| 095 | n/a (instrument or theory) | 9 core claims; evidence extracted from ~40 RESULT.md files (e.g. C1 11 support / 3 contradict; C5 and C8 rest on 2 rows each); 500 shuffle trials; prior sweep over 5 values; 34 self-tests | none | 1 | shuffled-evidence null (500 trials, max deviation 0.016) and prior sweep guard the ranking; direction/strength labels are analyst judgment; +-5 log-odds clamp | CPU stdlib, milliseconds | instrument (meta-analysis over campaign rows) | Aggregation tool, not new data; posteriors inherit every row's thinness and rows are correlated (single model), which the RESULT states explicitly; C5/C8 rest on <=2 rows. |
| 096 | n/a (instrument or theory) | toy sweep: V in {8,16,32,64,128,256,512} x T in {1000,4000,16000} (21 cells) + capacity sweep H in {32,64,128,256,512} + beta smoke-sweep; novel-filler accuracy per cell | toy mean-pooled AE per cell: d=24, H=128, beta=0.10, 400 epochs, batch 2048, T tokens | 2 | 2-seed sd per cell recorded in out/results.json (e.g. 0.019 at V=8,T=1000) but not reported in RESULT; no CI on V* estimates | CPU numpy, main sweep 808 s (no GPU) | toy model | Theory/toy row; its empirical content (V* flat in T and H, beta-controlled) rests on 2 seeds per cell with no reported spread, and absolute V* is calibrated to 006, not predicted. |
| 097 | n/a (instrument or theory) | toy: 60 fixed multisets x 200 permutations each = 12,000 pooled vectors (d=64, V=40, n=6 tokens, H=4); 5 encoder variants; supplementary H and mu sweeps | numpy logistic-regression probes on pooled vectors; toy single-layer attention encoder (no learned weights stated) | 1 (not stated) | none (tie-corrected AUC point estimates; positive control poscontrol certifies probe power) | CPU numpy via uv, ~30 s + two supplementary sweeps | toy model (later H3 follow-up on real SONAR reported in an UPDATE note) | Theory/toy row; note P3 missed and the rank-tracks-heads sub-clause was wrong; the H3 real-SONAR follow-up revised rank from ~4 to ~23 PCs. |
| 098 | n/a (instrument or theory) | arithmetic over measured campaign quantities (066/068 capacity 460 and C_D 358 bits, rate 4.5 bits/tok, 063 knee 2.14 rows, 068 knee 70-83 tok, anatomy 15.79 cw with n=150/bin and CI [11.54,17.68]); one new corpus stat from 30 cached reconstructed FLORES heads (14.65 cw/row, 2.30 tok/cw) | none | n/a | none new; predictions judged against 2x tolerance bands; no error propagation from input uncertainties; 15.79 CI inherited | CPU, ~1 s | other (theory over SONAR measurements) | Derivation row; the 'within 10%' ratios carry no propagated uncertainty and the content-word conversion rests on a 30-sentence sample (RESULT calls this the soft joint); P6 missed. |
| 099 | n/a (instrument or theory) | 5 demos baked from frozen JSON of rows 080/065/071/086/087; 21/21 decode strings verified byte-identical to source | none | n/a | none needed (artifact fidelity checks) | CPU only, no inference; headless Chromium render check | instrument (artifact, no data claim) | Explainer artifact over frozen results; inherits the source rows' sample sizes (e.g. 071 and 087 demos). |
| 100 | n/a (instrument or theory) | packaged stimuli: agent_patient 2000, genitive 1200, causal 2000, temporal 2000 (+093 v3 families); end-to-end run completed only on the 48-d mock encoder; SONAR reproduction via frozen reference cross_all [0.4655,0.4833] plus a partial cell-level match | none new (battery probes as shipped) | n/a | battery bootstrap CIs in the mock run (e.g. 0.496 [0.485,0.508]); frozen reference CI for SONAR | CPU-only; fresh full 1024-d SONAR sweep did NOT complete (exhausted 15 GB and 31 GB RAM) | instrument (benchmark package) | Build row; the one substantive caveat is that the flagship SONAR null was not freshly reproduced end-to-end (P1 and P2 scored FALSE, Brier 0.314), only matched against the frozen reference. |

# Sample-size and compute audit of the consolidation and September follow-ups

| id | flag | items behind the headline | training set | seeds | uncertainty | compute | substrate | why |
|---|---|---|---|---|---|---|---|---|
| H1 | thin | English: 2,000 agent/patient items (200 propositions x 5 constructions x 2 orders) + 1,200 genitive items; cross-lexical cell = held-out-lexicon subset (split size not stated in RESULT). Japanese: 1,500 items per lexicon (750 propositions x 2 orders; 500 train / 250 test props), 3 disjoint lexicons; 3rd lexicon = 20 nouns / 12 verbs with 8 test nouns / 5 test verbs (S1 cell = 500 test items). Power cert: planted d=1.0, AUC 0.973 CI-lo 0.876 | English: the training-lexicon share of the 2,000/1,200 items (not stated numerically); Japanese: 500 train propositions x 2 = 1,000 items per lexicon (12 nouns / 7 verbs) | seeds 0-4 (5) | proposition-cluster bootstrap, n_boot 1000, 95% CI per cell; planted-signal power curve with CI; no permutation test | 4l box GPU phys3 (RTX A4000), real SONAR text_sonar_basic_encoder; wall time not stated; linear + MLP probes | real SONAR encoder (eng and jpn_Jpan), templated two-noun transitive stimuli (stimuli_v2 English; 061/062 Japanese case-marked SOV/OSV) | English null is well powered (2,000 items, 5 seeds, clustered CI); the Japanese retraction rests on a single third lexicon (8 test nouns, 250 test props) and a 3-lexicon comparison, with the Japanese power cert FAILing. |
| H2 | very thin | Natural bases per alpha cell: negation 40, tense 40, vertical 24 (104 sources; 1,196 texts, 1,424 decode jobs); negation 0.70 at alpha=1 = 28/40; rt_cos AUC 0.845 on 936 fluent/garbled decodes (597/339); tool flag: 229 high-alpha garbled, 90 alpha=1 fluent; 1,320 codex-judged records | 64 templated train pairs per operator per seed (3 operators x 2 seeds) | 2 operator seeds (0, 101); single decode run; single judge (local codex) | none (no CI/bootstrap on any rate; 8 Brier-scored predictions only) | 4l GPU phys0 (RTX A4000); decode 434 s, codex judge 1,242 s; SONAR-z encoder/decoder | real SONAR encoder+decoder; natural wnat sentences (6-26 words) steered by z + alpha*v; templated train pairs | Per-operator success/fluency rates rest on 40 (or 24) natural bases per dose with no interval; only the rt_cos AUC (936 decodes) has a moderate n; vertical later found 0/24 applicable (A5). |
| H3 | thin | Binding stimuli: 2,000 items (200 propositions x 5 families x AB/BA), per-layer mean-pooled states; fresh permutations: 40 fixed multisets x 6 words x 120 orderings = 4,800 sequences; rank estimate (22.8 PCs) from the 40-multiset permutation set | prop-clustered CV on the 2,000 items (fold sizes not stated); no separate train set for the permutation cells | not stated (single run) | none reported (proposition-clustered CV AUCs only; no bootstrap CI, no permutation test) | CPU only on 4l box; numpy/sklearn probes; reused binding_death states + fresh PrepoolEncoder extraction; wall not stated | real 24-layer SONAR encoder token states (L0-L24n); templated binding_death sentences and ungrammatical fixed-multiset word strings | No uncertainty reported anywhere; the ~23-PC and within-multiset 0.997 numbers rest on 40 multisets; E01 later showed the role-from-odd 0.85 is a noun fingerprint (0.51 noun-disjoint). |
| A1 | n/a | 5 overnight branches (meta-prediction over audit verdicts) | n/a | n/a | none (single Brier-scored forecast, p=.75, outcome FALSE, Brier 0.562) | none (CPU replays across A2-A5, gtr) | audit verdicts of direct_reader, slot_reader, sae, operator, gtr branches | A meta-forecast over five audit outcomes, not an experimental estimate. |
| A2 | thin | 200 test propositions on 32 disjoint nouns x 32 sources x 4 rows = 25,600 dependent rows per rendering (CONJ, SEP); 4 verbs / 6 verb pairs fixed; decoder witness 640 decodes (first 20 propositions) | 2,000 propositions (48 nouns) x 16 sources x 4 rows = 128,000 rows; 200-prop validation; N=500 learning-curve arm | 3 init seeds (0,1,2) per arm; 12 fits total | 2,000-draw multinomial bootstrap over 200 test propositions (seeds fixed); auditor added 32-noun cluster bootstrap (SEP CI widens to [.966,.992]) | original: 4l GPU1 (RTX A4000), 12 MLP fits 03:45-03:59 UTC; audit replay CPU only | real SONAR z of two-clause conjoined sentences + bare-word SONAR query embeddings; MLP 4096->3072->1 (12.6M params) | AUC is precise at 200 propositions, but the population units are 32 test nouns and a closed 4-verb bank (~333 train props per verb pair), and SEP is an in-distribution re-template (cos .90). |
| A3 | thin | 128 base groups x 4 cyclic rotations = 512 carriers, 2,048 present queries (oracle 2025/2048 correct); 6,144 present-query predictions replayed; absent control 512 (auditor: 8,192 pairs, 0 accepts); direct reader 4.9% on the same 2,048 | 300 four-fact carriers (frozen shared_context latents); direct/no-query ridge on 1,200 (carrier, query) rows | 1 data seed (train 9824, test 10824); deterministic closed-form ridge; no repeats | bootstrap over 128 base groups (seed 10824), 95% CI [.984,.993]; replicate count not stated | original on 4l (GPU not stated, run.log empty); audit replay CPU via ssh; ridge is seconds-scale | real SONAR z of 4-sentence count-grammar carriers (20 objects x 8 counts x 12 locations x 2 tenses); position-specific ridge heads + presence gate | 128 independent groups (<200) from one training draw on a closed single-template grammar; routed endpoint is identical to oracle (not independent), and the direct comparator is capacity-limited. |
| A4 | thin | 12 SAEs (2 families x 6 seeds) -> 15 seed pairs per family for count-weighted matched fraction (.132 vs .111); mass-weighted ~.46 recomputed by the auditor on 3 pairs (0_1, 2_3, 4_5) per family over 8,192 held-out c_pool rows; polarity utility .59-.63 from S3's 96 polarity clusters | 280k c_pool rows per SAE, width 8192, k=32, 40 epochs = 2,760 optimizer steps (unconverged, FVU .42-.44 still falling) | 6 per family (paired init and row order); auditor replay on 3 pairs | none formal (leave-one-seed-out ranges in lit-gap; all 6 per-seed paired differences negative, pair sd .002); no bootstrap on mass-weighted figures | original SAE training on 4l GPU2 (GPU model/wall not stated); audit replay CPU | SONAR c_pool (order-removed pooled state) + freshly trained per-sample TopK vs BatchTopK SAEs (8192, k32) on cpool text chunks | Unit is the seed (6 per family, one width/k/corpus, unconverged dictionaries); the ~.46 mass-weighted headline is an unregistered auditor replay on 3 pairs. |
| A5 / O1 | thin | 104 reused H2 natural sources (40 negation, 40 tense, 24 vertical) x 2 arms (alpha 0/1) = 208 judged rows; applicable 39/38/0; full success negation 15/39 (fable join; lit-gap SUMMARY says 14/39), tense 16/38; collateral not preserved 19/40 (47%) and 7/40 (17%, +17/40 uncertain); vertical 0/24 applicable | none new (H2 seed-0 directions reused; 64 templated train pairs per operator) | 1 operator seed; 1 blinded outcome rater (Codex); 2 source-applicability reviewers; alpha in {0,1} only | lit-gap: none by prereg ('no population CI'); fable audit added Wilson 95% over sources (negation full 15/39 [.25,.54]; collateral [.33,.63]) | none new (reused H2 decodes; CPU joins; aggregate_audit_v2 = 408 stdlib checks) | real SONAR z + H2 negation/tense/vertical directions on 104 purposefully mined wnat natural sentences, greedy decode; model-judged full-proposition rubric | 38-39 applicable sources per operator judged by a single model rater with Wilson intervals ~0.3 wide; vertical has 0/24 applicable (very thin, undefined estimand); 14 vs 15 negation count unresolved. |
| G1 | very thin | 64 count sources (unit = source) and 64 role strings = 16 proposition blocks x 4 forms (unit = 16 blocks); 4 arms = 512 outputs; clean20 count 60/64 [.875,.984], role 40/64 [.469,.781]; blind review of 128 clean20 outputs (0 false-correct) | none (third-party pretrained jxm/gtr__nq__32 hypothesiser + corrector, NQ-trained by Morris et al.) | 1 (greedy, beam 1, 20 correction steps; REPEAT re-decodes identical) | conditional group bootstrap 2,000 resamples (seed 90202; stated in PREREG.md, not RESULT.md), clustered by source / 16 role blocks | 4l (RTX A4000 per G3 RUNTIME.json; CUDA 12.4, torch 2.6.0); wall not stated | GTR-t5-base raw mean-pooled 768-d + vec2text inverter (pinned abe48a5); templated count grammar and S2-style role sentences | Role 40/64 rests on 16 blocks (CI .47-.78) and count on 64 sources, single deterministic run; 20/64 role outputs are unadjudicated parser-unknowns. |
| G2 | ok | 128 final sources per domain x 2 domains x 6 paired arms = 1,536 outputs (+64 calibration per domain: 62/64, 60/64); half-norm main 119->2/128, domain 124->0/128; blind calibration review 128; the paper's '3/128' is G3, not G2 | none | 1 (deterministic; perturbation seeds 90321/90322) | paired source-cluster bootstrap 2,000 draws (seed 90322): half-norm minus clean -91.4 pp [-96.1,-85.9] main, -96.9 [-99.2,-93.8] domain | GPU inference on 4l (GPU type not stated in g2/; final complete 04:03 UTC); audit CPU <1 s | GTR-base + vec2text; count grammar; arms clean, norm x0.5/x1.5, tangent noise cos .99/.94, counterfactual | 128 paired groups per domain with clustered CI and a -91 pp effect; but every perturbation arm is 88-100% parser-unknown (floor effect), so the endpoint cannot rank arms and the prereg's semantic review was not run. |
| G3 | ok | 128 fresh paired source groups, main domain only (clean 124/128, half-norm 3/128, counterfactual 121/128; effect -0.945); calibration 64/domain; the 'domain' vocabulary failed calibration and was skipped; model review 32 groups / 96 cases | none | 1 (deterministic) | preregistered one-sided Hoeffding bound (alpha .025) on a conditional finite frame: upper -0.705 < -0.20; no bootstrap | RTX A4000, CUDA 12.4, torch 2.6.0, vec2text 0.0.13 (RUNTIME.json); wall not stated | GTR-base + vec2text; fresh count-grammar combinations excluding all 2,920 historical tuples | 128 paired groups with a preregistered bound and a near-total effect; caveat is that only 1 of 2 registered domains ran and the parser-only core does not determine semantics. |
| E01 | ok | 8,000 pile-10k sentences / 226,392 tokens; test = 1,123 documents / 1,629 sentences / 45,402 tokens (content-token fraction .93); binding transfer on cached binding_death items (prop-level); random-init twin on the same corpus | 162,821 train tokens (+18,169 validation tokens for lambda selection) | all seeds fixed at 0 (corpus, split, P/Q, C6 init, bootstrap); single run | 1,000-replicate bootstrap over test documents (props for transfer), 95% CIs; paired bootstrap for trained-vs-random contrast | 4l GPU 2 (RTX A4000), 1,024 s total (extract 145 s, fit 830 s); fp64 ridge | real SONAR 24-layer encoder token states (L0-L24n) on natural pile text; ridge n-gram feature sets F0-F6; re-initialised twin encoder | Large token/document n with document-clustered bootstrap and an independent reimplementation; single data size and one seed (deterministic ridge). |
| E02 | thin | 2,000 binding_death items (5 families x 200 props x AB/BA); test per family = 50 propositions x 2 = 100 items on 22 held-out nouns (3 names), 74% common-common; objrel/nominal ranges over training-family sets; layer profile on the same items | 65 train nouns; D1 train families {active,passive} (per-family 150 props x 2 = 300 items; sizes vary by training set) | 3 fit seeds (0,1,2) but ALS converges identically (per-seed AUCs equal to 4 d.p.); random-init encoder seeds 0 and 1 | 95% bootstrap over the 50 test propositions per family (not nouns); replicate count not stated in RESULT | CPU only, 14,282 s (ALS bilinear readers); breaker replay CPU | real SONAR L24n z of templated binding_death sentences + bare-noun SONAR query embeddings; rank-64 bilinear reader | Every unseen-noun number uses the same 22 test nouns and 50 test props per family; the bootstrap resamples props not nouns, and the 3 seeds add nothing. |
| E03 | ok | E01's 8,000 sentences; test 1,123 documents / 1,629 sentences / 45,402 tokens (content tokens scored); random-init twin same split | 162,821 train tokens (E01 split); rank-64 G interaction fitted on train | seed 0 (partner derangement, C6 init); single run | 1,000-document test bootstrap 95% CIs (Q1 .035 [.034,.035]; Q4 .322 [.318,.325]) | 4l GPU 2 (RTX A4000), 126 s (states reused from E01) | real SONAR L24n token states on natural pile text; own-token + leave-one-out sentence-mean decomposition; random-init twin | Same large document-clustered n as E01 with tight CIs and a built-in positive control (random-init +.348); no separate breaker pass. |
| E04 | very thin | 37 common-noun test propositions x {active, passive, cleft} x 2 directions = 222 edits (19 test nouns); secondary 65 training propositions; parser audit 60 decodes; oracle .982, reader-dose .000/222, 2x .072 | E02 reader refit on D1 train (65 nouns, seed 0, r=64, lambda 1e3) | 1 reader seed (0); shuffled-label control seed 999; single run | noun-clustered bootstrap, 2,000 draws (resample 19 nouns; prop weight m_A*m_B), plus prop-level; 95% CIs (P3 .072 [.01,.23]; P6' .423 [.16,.77]) | 4l GPU 2 (RTX A4000), ~38 min | real SONAR encoder+decoder; templated binding_death propositions; additive edits along the E02 reader gradient and subspace projections; regex parser | 37 propositions / 19 nouns give CIs like [.16,.77] on the positive-control arm, which is why neither registered branch was licensed; the exact 0/222 is robust but the dose grid is single-seed. |
| E04b | thin | 160 fresh propositions (40 new nouns, 8 new verbs) x {active, passive, cleft} x 2 directions = 960 edits; 92,160 decodes; 5 draws for random-subspace controls; parser audit 40 decodes; S-9 .140, controls .124/.110, isotropic .024, oracle .995 | E02 reader (65 nouns) refit; fresh reader AUC .977-.980 checked on the new stimuli | 1 reader seed; 5 random-subspace draws; stimulus seed 20260911; single run | noun-clustered bootstrap 2,000 draws (40 nouns, prop weight m_A*m_B) + verb-clustered (8 verbs) + prop-level; 95% CIs (S-9 .140 [.07,.23]; paired diff -.001 [-.04,.04]) | 4l GPU 2 (RTX A4000), 2,127 s for 92,160 decodes | real SONAR encoder+decoder; fresh templated propositions; coordinate swaps inside rank-9 reader vs energy-matched PC subspaces; regex parser | 160 propositions (<200) on 40 nouns / 8 verbs with clustered CIs; the paired null difference [-.04,.04] is precise enough for the negative claim but the design is single-seed. |
| audit replication | very thin | Exact replication: 1,411 source rows / 211 groups x 8 arms; 577 surrogate test rows / 85 held-out groups x 2 surrogate arms (12,442 decodes); 'gate accepted 0 surrogates' = 0/577 per arm; historical 41-49% = 38/78 and 34/82 valid reads (ridge reads from 28 groups); 'strict pairs 5 and 2' = 5 reads from 3 groups and 2 from 2 groups out of 577; 'blind review 5/33' = 33 flips inside the first 60 of a 375-item packet | no new fitting; historical SBERT->SONAR ridge surrogates (ridge 834 rows / 126 groups; small ridge 396 rows / 60 groups) | seed 0, single deterministic run; bootstrap seed 4242 | 10,000-replicate whole-group cluster bootstrap (85 groups), descriptive; 4 Brier-scored high-probability predictions (Brier 0.0025-0.04) | 4l GPU0 (RTX A4000 box), 382.6 s wall; fp16 decoder batch 32 beam 5 | real SONAR encoder+decoder; historical natural active/passive/edited/nominal sentences; SBERT-ridge surrogate latents; noise/interpolation arms | Replication and 0/577 acceptance are solid (85 groups, clustered bootstrap), but the quoted strict-pair counts come from 3 and 2 groups and the 5/33 review is a 60-item convenience subset with no interval. |
| count_capability | thin | 64 fresh source clusters x 3 fact counts x 2 arms = 384 decodes; each headline numerator (60/25/0) out of 64; beam ablation = same 64 latents x 3 beams (1,152 decodes): clean 60/59/59, 25/20/21, 0/0/0 | none (no fitting) | data seed 1733; single greedy run (beam-1 repeat reproduced all 384); bootstrap seeds 2733/3733 | paired 64-cluster bootstrap, 10,000 replicates: 1->2 facts -54.7 pp [-67.2,-42.2]; 1->4 -93.75 [-98.4,-87.5]; beam5-beam1 two-fact -7.8 pp [-18.75,+4.69]; blind reviews of 40 and 29 items | 4l GPU0; float32 SONAR greedy batch 8; capability run 47 s; beam runs 36-44 s; GPU model not stated in-unit | real SONAR encoder+decoder; templated count grammar with 1/2/4 nested facts | 64 source clusters only; the 60->0 collapse is robust in sign, but the beam CI on two-fact spans zero (beam 5 nominally lowered 25->20), so 'beam does not move it' is unresolved at this n. |
| count_monitor | ok | Main: 300 fresh source clusters x 7 arms = 2,100 outputs (.85 self gate accepts 1,582 incl. 379 mismatches); domain 200 x 7 = 1,400; calibration 200 x 7 = 1,400; 50%-coverage rows = top 1,050 main / 700 domain; source-aware 0/1,050 vs self 280/1,050 | 200 calibration clusters (1,400 outputs) used only to pick the .99 self threshold; no learned model | data seeds 54101-54103; decode seed 8821; single run; bootstrap seed 821 | 10,000-replicate source-cluster bootstrap, all 7 arms paired (pessimistic risk 1.34-4.31%); Clopper-Pearson sensitivity bounds; 108 independent recomputed checks; 20-item blind review | 4l GPU0; float32 SONAR greedy batch 8; test split 76 s; GPU model not stated in-unit | real SONAR encoder+decoder; one-fact count grammar; arms clean, noise .94/.85/.70, counterfactual latent, norm x.5/x1.5 | 300 independent clusters with clustered bootstrap; but 278 of the 379 mismatches are from the deliberately counterfactual arm, and the 0/1,050 is a rank-selected mask with a degenerate CI. |
| count_continuous | thin | 1,400 old calibration outputs (200 source clusters x 7 arms), no new outputs; source-aware .99146 accepts 705/1,400 (0 mismatch, upper risk 5.882%); latent-only: no candidate meets the rule at any coverage | the 200-cluster calibration set is itself the fitting set (exhaustive threshold search); the 1,000-source final cohort was cancelled | no decoding; bootstrap seed 4821 | 10,000 source-cluster bootstrap count tables; 95th-percentile upper limit on pessimistic risk; Brier P1 .7225 (failed), P2 .0225 | CPU analysis on existing outputs; not stated | reuse of count_monitor calibration outputs (real SONAR, one-fact count grammar) | Calibration-only search on 200 sources with no fresh test outputs; the source-aware miss is by 0.9 pp on the bootstrap upper limit. |
| count_targeted | thin | 128 fresh source tuples x 13 arms = 1,664 outputs; at cosine .99: 137 parsed mismatches across 384 targeted outputs (3 directions x 128) vs 0/128 isotropic; 128 paired units | no new fitting; attacker uses frozen count_reader ridge (200 calibration clusters / 400 rows) | data seed 8823; one seeded isotropic direction per source; single decode run | paired 128-source bootstrap, 10,000 replicates: targeted minus isotropic 35.68 pp [33.33,38.02]; blind 57-pair review | 4l GPU0; float32 SONAR greedy batch 8; wall not stated; GPU model not stated in-unit | real SONAR encoder+decoder; one-fact count grammar; reader-guided vs Gaussian isotropic norm-preserving perturbations at matched cosine | 128 sources (<200) with a single random isotropic direction per source; the effect is large and clustered-bootstrapped, but the isotropic '0' baseline rests on one draw per source. |
| count_norm | ok | 300 main + 300 domain fresh sources x 8 arms = 4,800 outputs; main half-norm 154/300 -> restored 283/300 (+43.0 pp); 96 location-only mismatches -> 0 (domain 31 -> 0); monitor adaptation 2,100 outputs per cohort | no model fitting; target norm = median of 200 old calibration embeddings; .99605 threshold frozen from old calibration | decode seed 8821; single run; bootstrap seed 6821 | paired 10,000 source-bootstrap 95% CI (+43.0 pp [36.3,49.7]; domain +32.3 [26.7,38.0]); iid-group sensitivity bounds; blind 96-pair review | 4l GPU0; float32 SONAR greedy batch 8; test split 85 s; GPU model not stated in-unit | real SONAR encoder+decoder; one-fact count grammar; radial scaling x.5 then restoration to median calibration norm | 300 independent sources per cohort, paired clustered bootstrap, replicated on a second vocabulary; the 96->0 count is exact. |
| count_reader | ok | 300 fresh main sources per field (+300 alternates = 1,200 clean/counterfactual predictions); 300 domain sources for count/tense; rescue subset = the 96 main sources with half-norm location-only decoder errors (96/96) | 200 old calibration clusters x (original + alternate) = 400 rows; 5-fold CV 400/400 | seed 7821; single fit, alpha .01; label-shuffle control seed 17821 | 10,000 paired source-cluster bootstrap (degenerate [1,1] at 300/300); one-sided binomial lower bounds 99.01% (300/300) and 96.93% (96/96); shuffle control (e.g. 8/300 object, 151/300 tense) | CPU ridge on saved embeddings; wall not stated | real SONAR embeddings (L2-normalised), centred ridge with intercept, closed class inventories; templates and main vocabulary recur between train and test | 300 unseen source tuples per field with an exact result and a shuffle control; the 96/96 rescue is a decoder-selected subset under 100 but also an exact count. |
| historical_norm | very thin | 577 held-out rows / 85 groups x 6 conditions = 3,462 outputs; strict-pair valid reads ridge 5 -> 3, small ridge 2 -> 5 (each within ~3 groups); legacy valid 78 -> 59, 82 -> 79; gate accepted 0 -> 0 | no new fitting; normalisation target = median clean norm of 126 / 60 original training groups | single run (raw surrogate decodes reproduced all 577 prior outputs); bootstrap seed 8822 | paired 85-group bootstrap, 10,000 replicates (ridge strict -0.35 pp [-1.40,+0.51]); blind 117-pair review | 4l GPU0; fp16 decoder batch 32 beam 5; 101 s wall | real SONAR (fp16 decoder); historical SBERT->SONAR ridge surrogate latents; median-norm rescaling | The strict-pair change is 5->3 and 2->5 reads inside about three groups, which the RESULT itself calls too sparse; 0->0 acceptance across 577 rows is exact. |
| multifact_reader | thin | 150 test source groups with nested 1/2/4-fact contexts (450 carriers); present queries 150/300/600; 2,850 absent queries per length; joint accuracy 103/150, 29/300, 11/600; positive-control subset 52 two-fact carriers (10/104), 2 four-fact (0/8) | 900 carriers over 1/2/4 facts; per-object ridge readers see 77-131 carriers each (20 objects); 5 CV shards | data seed 8824; single fit, alpha .01 | paired 150-source bootstrap, 10,000 replicates (1->4 facts -66 pp [-73.3,-58.7]); independent replay; blind review of two decodes | CPU ridge + 450 decodes on 4l GPU0 (float32 greedy batch 8); wall not stated | real SONAR whole-text embeddings; multi-fact count grammar; per-object ridge readers | 150 source groups and per-object readers trained on ~100 examples each; the RESULT itself says the reader is too weak at one fact to interpret; the 0/8 control subset is very thin. |
| shared_context | thin | 150 test source groups reused across all 9 train x test context cells (1,800 predictions); calibration screen 100 sources (100/100); decoder 450 carriers (120/150, 21/150, 0/150 verified) | 300 sources per training context (3 contexts share the same 300 first-target facts); 100-source calibration | data seed 9824; single fit | 10,000 paired 150-source bootstrap (four-fact context gain +98 pp [95.3,100]); iid binomial sensitivity (150/150 -> lower .980) | 4l GPU0 float32 greedy batch 8 for 450 decodes; ridge on CPU; wall not stated | real SONAR embeddings; multi-fact count grammar; shared ridge readers; second vocabulary in train and test; queries always first-position fact | 150 groups (<200) reused across every cell; effects are extreme (100% vs ~0-7%) so sign is secure, but intervals are degenerate and templates recur. |
| lexical_fidelity | thin | 128 OLD calibration source vectors (64 original + 64 fresh; each = 32 object-anchor pairs x 2 orientations) x 3 length penalties = 384 beam-5 decodes; wins/losses over a frozen 521-candidate likelihood table per cohort of 64: penalty 1: 0/17/47 and 0/18/46; penalty 0: 1/17/46 and 0/18/46; penalty 2: 0/17/47 and 1/18/45 | none (no fitting); the bank is reused development/calibration data | single beam-5 decode (first arm reproduced all 128 historical outputs); bootstrap seed 90531 | 10,000 paired cluster bootstrap over 32 object-anchor groups per cohort for accuracy differences only (pen0-pen1 -1.6 pp [-5.5,+1.6]); no interval on wins/losses; 34-output blind review | 4l GPU3; float32 SONAR beam 5 batch 8; COMPLETE stamped 12 s after start (model load likely excluded); GPU model not stated in-unit | real SONAR decoder on frozen O2 calibration embeddings (spatial-relation grammar); candidates rescored under the native generation objective | 64 sources per cohort (32 pairs) on a reused development bank, no uncertainty on the wins/losses counts, and no RESULT.md exists (numbers only in development/out/SUMMARY.json and STATUS.md). |
| binding/iterations | ok | Base run: 1,500 held-out test propositions x 2 directions x 2 unseen constructions = 6,000 rows on 24 test nouns / 6 verbs. iteration2 (source of order 0.959 vs role 0.526): 1,500 fresh-lexicon props (24 new nouns x 12 new verbs) x 2 x 2 = 6,000 rows; decoder 4,800 sentences. iteration3: 1,500 props (32 nouns x 12 verbs) x 2 x 6 constructions = 18,000 rows; decoder 4,800 | 4,000 propositions x 2 directions x 2 constructions = 16,000 rows (48 nouns / 12 verbs) in every sub-run | MLPs: 3 init seeds (0,1,2) per sub-run; ridge readers deterministic | paired proposition-cluster bootstrap, 2,000 draws, fixed readers; iteration2/3 add crossed noun x predicate bootstrap (2,000) and noun delete-one ranges; no permutation test | 4l GPU1 (RTX A4000), SONAR float32; base ~70 min end-to-end (48,800 encodes); iteration2: 24,836 encodes, 25 MLP fits (1.4-11.7 s each), 4,800 decodes; iteration3: 30,044 encodes, 24 fits; wall for iter2/3 not stated | native SONAR encoder+decoder; templated two-noun transitive sentences (profession nouns) in active/passive/cleft/objrel (+subjrel/objcleft); MLP readers on [z, bare-word queries] | 1,500 proposition clusters, 3 seeds, clustered bootstrap plus crossed-lexicon sensitivity; but only 24-32 nouns x 12 verbs, so crossed intervals are wide (role lift [-0.03,+0.08]) and the 0.526 hides cleft 0.754 / objrel 0.298. |
| binding/multievent | very thin | Candidate-likelihood reader: 200 base propositions -> 3,200 sources -> 9,600 query rows (pooled AUC 0.725). Decoder witness (paper's 242/320): 320 greedy decodes = first 20 base propositions x 16 cells; 122 accepted opposite-role sources; candidate reader correct on 5/122 = 4.1%. direct_multievent (no RESULT.md): 200 base props -> 6,400 sources -> 25,600 rows; decoder 640 | multievent: none (frozen native-decoder likelihood reader); direct_multievent: N=500 and N=2,000 base propositions (32k / 128k rows), 200-prop validation | multievent: no fitting, 1 run; direct_multievent: 3 fit seeds per arm (12 fits) | paired bootstrap 2,000 draws over the 200 base propositions for the AUC; decoder counts and the 4.1% have no interval ('descriptive and conditional on parser acceptance') | 4l GPU1 (RTX A4000); multievent scoring ~6 min (80,060 teacher-forced candidates, 320 decodes); direct_multievent 41,600 sources / 166,400 queries encoded, 12 fits, 640 decodes; wall not stated | native SONAR encoder + greedy decoder (beam 1); two-clause conjoined sentences from iteration3's 32 fresh nouns and a fixed 4-verb bank | The paper's 242/320 and 4.1% rest on 20 base propositions (RESULT: 'based on 20 base propositions, not 320 independent populations') with no interval; the 0.725 pooled AUC collapses to ~0.51 within clause-order strata. |
| binding/predicate_holdout | very thin | NO RESULT.md; only development stages ran. development_v2 (source of 0.426 and 0.998): strict SEEN-predicate control panel = 9,312 sources / 33,600 queries over 12 training verb families with train nouns; the only held-out-family panel scored (validation) = 3,200 sources / 11,520 queries over 4 unseen families (two-event joint 0.061/0.080/0.166, single AUC ~0.993). The registered 8-family final (33,600 sources) was never encoded. paired_objective_v1 (source of 0.077): same panels, 9 fits | 21,600 sources / 81,600 queries from 24 training verbs (12 families) and 47 train nouns | 3 fit seeds per arm (0.426 = mean of 0.175 / 0.514 / 0.587) | none by prereg ('no query-independent error bars', 'no p-values or CIs'); seed-within-family then family macro means only | 4l GPU3, legacy venv (torch 2.6); dev_v2: 34,480 encodes + 6 fits, 486.9 s; paired_v1: 9 fits on cached embeddings, 621.6 s; final never run | native SONAR encoder; single-event and two-event conjoined sentences over 24 two-verb families; MLP 4096->3072->1 on [z, nounA, nounB, verb] | Three seeds ranging 0.17-0.59 with no interval; the quoted 0.426 is the seen-predicate control panel, not a held-out-predicate result; the registered held-out final was never run. |
| binding/direct_confirmation | n/a | NO RESULT.md; no scientific result. A fresh cohort of 400 base propositions (12,800 sources / 51,200 queries) was prepared but never encoded or scored; sub-stage calibrations on 64 historical sentences / 256 query rows from 2 base propositions (reader_calibration 91/91 PASS; encoder_calibration_v2 FAIL 15/80; attention/precision diagnostics FAIL 41/246, 38/246) | none (9 frozen N2000 direct_multievent checkpoints reused) | 3 frozen seeds x 3 arms reused; no new fits | none (deterministic tolerance checks) | 4l GPU3; each diagnostic ~20 s to a few minutes | native SONAR encoder numerical reproducibility (batch regrouping / TF32 / attention kernel) on direct_multievent readers | No estimate produced; the thread ended at a failed encoder-compatibility gate (relative-L2 drift 6.8e-4 > 1e-4). |
| binding/lowdose | thin | Reused iteration1 test rows: 1,500 propositions x 2 orientations x 4 constructions (3,000 per construction); 24 nouns / 6 verbs | 4,000 propositions (reused iteration1 vectors, train-only PCA 64/32); 30 ridge fits | deterministic ridge; plant seeds 91004/91003 | proposition bootstrap 2,000 draws + crossed noun/predicate sensitivity: bilinear minus concat 0.406 [0.385,0.427] proposition, [0.227,0.595] crossed | CPU ridge on cached vectors; wall not stated | planted synthetic global / query-addressed codes on projected real SONAR backgrounds; ridge z-only / concat / bilinear readers | Reused rows, no independent cohort, only 24 nouns x 6 verbs so the crossed interval is very wide; RESULT calls it 'not a new independent confirmation'. |
| binding/query_order | ok | Reused iteration3 fresh test: 1,500 propositions x 2 orientations x 6 families x 2 query orders | 4,000 propositions; 32,000 optimizer rows per arm; 12 fits | 3 fit seeds (0,1,2) per cell | paired proposition bootstrap 2,000 draws: AP gain +0.166 [0.163,0.170] | 4l GPU1; wall not stated | native SONAR vectors; MLP width 3072 on [z, qA, qB]; iteration3 template sentences | 1,500 clusters, 3 seeds, clustered bootstrap; reused cohort ('fitting replicates, not independent data populations'). |
| binding/semantic_reader | very thin | 160 packet cases: 64 parser-unknown decoded outputs spanning 46 source propositions (17/64 resolved), 32 clean references, 64 constructed controls | none (pretrained assistant reader) | 1 packet (sampling seed 91041) | source-proposition cluster bootstrap 2,000 draws on coverage [15.5%,39.3%]; accuracy bootstrap degenerate at 100%; iid one-sided 95% error bound ~16.2% for 17 successes | no GPU; reused decoded outputs; model-assistant reading | reused iteration3 SONAR greedy-decoded texts; assistant reader given decoded text + queried identities + predicate | Headline rests on 17 resolved cases from 64 unknowns (46 propositions); zero observed errors gives no usable accuracy bound. |
| binding/template_reader | ok | Reused iteration3 fresh test: 1,500 propositions x 2 orientations x 4 assessment constructions | none (fixed cosine to canonical candidate encodings) | n/a (deterministic) | paired proposition bootstrap 2,000 draws: macro 0.674 [0.669,0.678]; correct-predicate premium +0.039 [0.034,0.045] | 4l GPU1; 6,348 candidate texts encoded batch 128; wall not stated | native SONAR encoder; cosine of source z to canonical active/passive candidate embeddings | 1,500 clusters with clustered bootstrap and no fitting; reused cohort is a diagnostic, not fresh confirmation. |
| binding/decoder_likelihood | thin | Reused iteration3 decoder cohort: 4,800 source rows = 200 propositions per domain x 2 domains x 2 orientations x 6 families; 3,148 unique candidates, 60,748 candidate sequences scored | none (frozen SONAR decoder teacher-forced likelihood) | n/a (deterministic); shuffle seed 91072 | paired proposition bootstrap 2,000 draws: macro 0.753 [0.730,0.774]; premium +0.053 [0.029,0.076] | 4l GPU1; scoring 219.9 s; two analysis reruns after bugs | native SONAR decoder log-likelihood of canonical active/passive candidates given source z; iteration3 template sentences | Only 200 fresh propositions per domain (400 total) on a reused cohort; bootstrapped but under ~200 fresh clusters for the primary fresh-domain number. |
| U1 | ok | 1,024 held-out test master chains (16 sentences each), all slots scored at depths 2/4/8/16; 128-candidate retrieval per target; '92%' = depth-16 mean-baseline raw cosine 0.3142 / ridge 0.3414 | 2,048 training chains (nested N=256 arm); 512 calibration chains for alpha | 1 deterministic ridge fit per cell; one shuffled-label permutation (seed 90103); no repeats | percentile bootstrap 2,000 resamples of the 1,024 whole test chains, paired: N-effect +0.071 [0.070,0.072]; depth effect -0.303 [-0.307,-0.299] | 4l GPU2 (RTX A4000), SONAR float32 batch 32; 57,344 constituents + 4x3,584 chains encoded in ~3 min; ridge seconds; torch 2.9.1 / sonar-space 0.5.0 | native SONAR text_sonar_basic_encoder (NOT a Llama TAE; the llama-3b-split file is only a text corpus); artificial cross-document chains of 25-90-char sentences; ridge chain->constituent | 1,024 document-disjoint test chains with a 2,000-draw chain-clustered bootstrap; single fit is deterministic ridge. |
| U1B | ok | Same 1,024 U1 test chains, posthoc rescoring; bank of 16 own-chain + 112 outside candidates; depth-16/N2048 ridge retrieval 0.0284 raw -> 0.2603 centred | none new (saved U1 predictions; training slot means reused) | none (deterministic; distractor seed 90106) | paired bootstrap 2,000 resamples of 1,024 chains: +23.19 pp [22.58,23.79] | CPU/GPU2, ~30 s | same SONAR artificial-chain cohort and frozen U1 ridge predictions | 1,024 chains with clustered bootstrap; explicitly a posthoc diagnostic on the same cohort. |
| U2 | ok | 512 of the 1,024 U1 test chains x 16 cyclic rotations = 8,192 rotated chains, every constituent at every position; oracle-boundary span mean-pool retrieval 0.997, first4-last4 difference 0.010 | none new (U1 N2048 ridge refit, verified to 1e-8) | none (deterministic; bootstrap seed 90141) | bootstrap 2,000 resamples of 512 whole chains keeping shifts paired: local minus ridge last4 0.979 [0.977,0.981] | 4l GPU2; 16 shift forwards with token-state capture, ~7 min | native SONAR encoder token states; oracle sentence-span mean pooling vs global pooled vector; artificial chains | 512 chains x 16 positions with chain-clustered bootstrap; local pooling uses oracle boundaries and uncompressed token states, so it is not a fixed-vector competitor. |
| U3 | ok | 1,024 U1 test chains, depth 16, all 16 slots (primary last 4), K128 centred retrieval | N=256 and N=2,048 chains; 512 calibration chains for alpha/gamma | deterministic kernel ridge; shuffled permutation seed 90210+N | bootstrap 2,000 resamples of 1,024 chains: RBF-linear last4 -0.0002 [-0.0022,0.0017] | 4l GPU2 float64 eigensolves; ~70 s | same SONAR artificial-chain vectors; linear vs Gaussian-RBF kernel regression | 1,024 chains with clustered bootstrap; reused adaptive cohort. |
| U4 | ok | 256 disjoint 64-source pools (4 consecutive U1 test chains each) at depths 2-64 for the additive oracle; native SONAR comparison at depths 2-16 on the first chain of each pool | rho from 32,768 training constituent vectors; nothing fitted to test | none (deterministic; bootstrap seed 90401) | bootstrap 2,000 resamples of the 256 pools; raw formula MAE across six depth means 0.0012 | CPU, 2 BLAS threads; seconds; no new encoding | reused SONAR constituent and chain vectors; unit-normalised additive codebook oracle vs native pooling | Geometric arithmetic check on 256 pools with bootstrap; makes no capacity claim. |
| HISTORICAL_SOURCE | very thin | Source-only reconstruction of the historical natural-chain depth curve: at depth 16, 231 windows from 16 documents / 27 paragraphs; 196 train / 35 test windows from 3 documents, 11 sharing a training paragraph; 110/560 test constituents also in training; depth 2: 8,000 windows | historical: 196 windows at depth 16 (vs 1,024 input dims); 6,800 at depths 2-5 | n/a (deterministic counting) | none (exact counts; no cosine recomputed) | CPU only, no model calls | text corpus nickypro/llama-3b-split used as SOURCE TEXT for the historical tae_comp_unbind3.py SONAR chains; no Llama model involved | The audited historical depth-16 point rests on 35 test windows from 3 documents with train/test paragraph overlap and 196 training examples; the audit itself is exact but exposes this. |
| S1 | ok | 8 frozen checkpoints (widths 8192-65536 x seeds 0/1, k=32) = a census; per checkpoint 8,192 held-out c_pool rows for regrouping, 4,096 for threshold calibration, 2,000 vocabulary-disjoint paraphrase pairs for invariant masks, first 60,000 rows for L1 mass. NB the paper's '45-92%' are the R16 pilot's figures (one checkpoint, 2,048 rows); S1 proper gives 50.2-56.0% shuffled / 91.2-94.5% norm-sorted | none new (threshold on 4,096 rows); underlying SAEs: 280k c_pool rows, 40 epochs, batch 4096, AuxK 256 | 2 training seeds per width (historical); deterministic inference, 1 run | none; prereg declines any CI ('census of eight checkpoints'); 4 registered predictions held | RTX A4000 (4l GPU2/3), torch 2.9.1+cu128; 4-10 s per checkpoint; width 65536 needed a memory-repair rerun | SONAR c_pool (order-removed pooled 1024-d state) + historical W40 BatchTopK+AuxK SAEs k=32, widths 8192-65536; per-sample vs batch/shuffled/norm-sorted vs calibrated global threshold | 8,192 rows x 8 checkpoints with a 50-94% effect is unambiguous; only 2 seeds, no CI by design, and 'threshold identical' is approximate (1/8192 rows flipped on h32768_s0). |
| S2 | thin | 6 paired training seeds per family -> 15 dependent seed pairs per family; count-weighted matched fraction sample .1320 vs batch .1114; activation correlation on 8,192 held-out rows; 'mass-weighted ~0.46' is NOT in any lit-gap SUMMARY (fable audit replay on 3 pairs); 'polarity .59-.63' is the S3 hard contrast (96 clusters) | 12 new SAEs (2 families x 6 seeds), width 8192, k=32; 280k c_pool rows, 2,760 optimizer steps (unconverged); 0 dead features | 6 per family (identical init + row order across families) | leave-one-seed-out ranges only ('sensitivity not CI'); no bootstrap; all 6 per-seed paired differences negative (-.017 to -.024); P1 and P3 failed | 4l box, CUDA_VISIBLE_DEVICES=2; GPU type and wall time not stated in sae_training files | SONAR c_pool + freshly trained per-sample-TopK vs BatchTopK SAEs (8192, k32), evaluated under per-sample and threshold inference | One width, one k, one corpus, 6 seeds, unconverged dictionaries; the 0.46 mass-weighted headline rests on 3 pairs replayed by the auditor, not a registered analysis. |
| S3 | thin | Bank of 288 clusters (96 role / 96 polarity / 96 count) x 4 strings = 1,152 strings in 16 lexical blocks (only exchangeable unit); 8 checkpoints x 2 rules = 16 cells, invariant mask (985-1,414 features) vs 10 matched control masks; premiums polarity +0.3998, count +0.1584, role -0.0934 | none (frozen masks from 2,000 paraphrase pairs; controls matched on 60k activation descriptors) | 2 training seeds x 4 widths; 10 control draws; no repeat of the bank | 16-block paired bootstrap, 2,000 resamples (pooled premium [.144,.166] per-sample); per-family CIs not in RESULT.md; conditional on fixed bank/models | not stated (encoding on 4l; GPU type/time not in RESULT/FREEZE) | SONAR c_pool + 8 W40 BatchTopK SAEs; controlled template bank (active/passive for role and polarity, existential/fronted for count); baselines native z, c_pool, bag | Only 16 lexical blocks; the fable audit shows the role result is a bank artefact (96/96 role pairs are pure order swaps on an order-removed input) and count is 1.0 for every baseline; only polarity is a real signal. |
| O2 | n/a | '17/64': operator_transfer calibration, 32 above/below pairs = 64 clean strings (gate 58). '55/64': calibration_v2, 32 fresh filler pairs = 64 strings. Search: 128 saved calibration vectors x beams 1/5/10 (fresh 55/58/58). '152 final sources never decoded': 128 physical (64 pairs x 2) + 24 polysemy strings, never encoded. '512-source frame bank at 84.6%': 64 unused filler pairs x 4 frames x 2 directions = 512 strings; beam1 433/512 = 84.57% [.764,.916] (preregistered primary beam5 = 421/512 = 82.23%); all 79 mismatches are object/anchor nouns ('clock' 25/40 wrong); blind sample 256 labelled. Likelihood_v2: 521 candidates x 2 = 1,042 scores on 128 vectors | operator directions v0/v1 reused from H2 (64 above/below train pairs per seed); no refit; frame bank has no fitting | 2 operator seeds (0, 101; v1 has cos .995 with v0); decoding deterministic; 1 run per stage | search: 95% group bootstrap over 32 groups; frames: 10,000 paired group bootstrap over 64 filler groups (descriptive); likelihood: none; calibration gates: none; the 4 original final-editing forecasts are unscored | SONAR encoder (3.06 GB) + decoder (4.52 GB) float32 on 4l (GPU type not stated); frames run 17.5 s / 21.8 s, peak 3.7 GB; search 3-5 s per beam | SONAR z with the H2 vertical direction (never applied to the final bank); controlled physical-relation templates ('The X is above/below the Y', 4 frames); narrow regex parser; beam 1/5/10 | The editing claim was never run (unrun tier); the calibration numbers 17/64 and 55/64 rest on 32 pairs each (very thin) and the 84.6% frame bank on 64 filler groups over a 16-object x 12-anchor vocabulary (thin), and 84.6% is the beam-1 control arm rather than the preregistered beam-5 primary (82.2%). |
| R01 pilot | ok | 2,000 synthetic test examples per seed x load; 5 seeds x 4 distractor loads (0/2/6/14); headline row is load 0, mean over 5 seeds (global linear .507, query-bilinear .999) | 4,000 synthetic examples per seed x load | seeds 0-4 | percentile bootstrap over the 5 seed-level AUCs, 10,000 draws; shuffled-label bilinear control (~.48-.51); RESULT calls intervals exploratory | CPU only on 4l, 1.10 s; dimension 16; NumPy/scikit-learn | fully synthetic: Gaussian fillers, two orthogonal role matrices (planted rotation code) vs additive shift code; NOT SONAR | 5 seeds x 2,000 items give tight seed-level intervals for an instrument-mismatch demonstration; the caveat is scope (d=16 synthetic code), not n. |
| R06 pilot | thin | 64 items = 32 counterfactual pairs, 4 families x 16, only 48 unique sentences and 8 lexical frames; 5 arms = 320 decodes + 64 ablation + 8 baseline = 392 decodes | none (no fitting) | seed 0, single greedy run; one seeded shuffle permutation; bootstrap seed 1 | 10,000-replicate frame cluster bootstrap over 8 clusters, later corrected to stratified 8 action + 8 number frames; premium 0.953 [0.891,1.0] unchanged; perfect-vs-zero arms give degenerate intervals | 4l GPU0 RTX A4000, 24.9 s including model load; SONAR float32 greedy batch 8 | real SONAR encoder+decoder; hand-authored templated sentences (role reversal, negation, tense, number) | 64 items but only 48 unique sentences and 8 lexical-frame clusters on a single seed; the 64/64 vs 0/64 exact counts are a robust positive control, while the premium CI lower bound rests on 8 frames. |
| R06/R09 | thin | Same 64 items (48 unique texts, 8+8 frames) x 2 noise seeds = 128 per noise level (not 128 independent sentences); clean 64, counterfactual 64; 512 decodes + 512 re-encodes; .94: 127/128 correct, 128/128 accepted; .8: 106/128 correct, 0/128 accepted; counterfactual 64/64 accepted | none; .85 gate frozen from H2, not fitted | 2 independent Gaussian noise seeds; decoding deterministic | stratified 10,000-replicate bootstrap (8 action + 8 number frames) superseding the original joint-frame bootstrap: accuracy change at .8 = -17.19 pp [-23.44,-10.94]; acceptance changes have degenerate intervals | 4l GPU0 (RTX A4000), 32.75 s including load and analysis | real SONAR encoder+decoder; templated sentences; isotropic Gaussian orthogonal noise at fixed cosine .94/.8/.6, norm-preserving | 64 sentences (48 unique) x 2 seeds with 16 frame clusters; the all-or-none acceptance results are exact and geometrically expected, so fragility is low, but the paper lists this at T3 rather than pilot. |
| R16 pilot | thin | 2,048 held-out c_pool rows from the historical 20,000 validation pool; threshold calibrated on a disjoint 4,096; 44.97% = 921/2,048 rows with any support change under shuffled batching (mean Jaccard 0.9765); norm-sorted 91.70% | SAE pre-trained (h16384, k32, BatchTopK, seed 0, 300k rows); fixed threshold fitted on 4,096 calibration rows to mean L0 = 32 | one SAE seed (0) and one split seed (20260909); single run; not formally preregistered | none on the 45% fraction; paired bootstrap CI only for the FVU delta (replicate count not stated); provenance rerun reproduced aggregates exactly | 4l GPU2 RTX A4000, 3.04 s; torch 2.9.1 | retained SONAR c_pool SAE (h16384); pooled-state embeddings of held-out validation-pool texts | 2,048 items is adequate and batch dependence is essentially a deterministic property of the rule, but only one SAE seed/width was tested and no interval is reported for the headline fraction. |
| multilingual pilot | very thin | 24 assistant-authored FR/EN/ZH triples (10 contrast pairs + 2 tu/vous pairs); 9 routes x 24 = 216 generations + 72 pivot generations (289 outputs per decoding mode incl. repeat check); greedy and beam-5 runs; direct = via-English 13/24 (greedy), 14/24 (beam 5); 20 ordinary contrast items ranked | none (no fitting) | seed 0; single run per decoding mode; one repeat item as reproducibility check | none (exact-match counts and mean cosines; no CI; assistant inspection, no blinded bilingual scoring) | one NVIDIA RTX A4000 (4l), 29.7 s (greedy) and 44.9 s (beam 5) including model load; SONAR float32 batch 8 | real SONAR text_sonar_basic_encoder/decoder (fra, eng, zho); hand-authored short sentences | 24 hand-authored items from 10 contrast families, single seed, no interval, exact-match agreement is a route-consistency diagnostic that the RESULT itself says is not translation accuracy. |
