# RESULT — 044 splitting-absorption (harvested 2026-08-03; run ALL DONE 02:06:07Z, total ~2 min on box)

**Tier: T3-exploratory.** Full chain ran: 2× h2048 e40 trains (48 s each, 2 GPUs) → decoder-direction
matching on 4 adjacent width pairs + cross-seed matching at 5 widths (28 s) → FVU + 043 probe battery
at 5 widths (21 s). Artifacts in `out/{match_results_full.json, probe_results_full.json,
atom_match_stats_full.npz, h2048_s{0,1}_train.jsonl}`. Nothing truncated; smoke ran first and passed.

## Gates

| Gate | Outcome |
|---|---|
| G1 instrument identity | **PASS** — all 10 ckpts reproduce logged cpool-val FVU within ±0.012 (≤ ±0.02 band); new h2048: val_fvu 0.434/0.435 ≪ 0.7, dead 0.0% |
| G2 matcher certification | **PASS** — self-match min cos 0.999996 (100%); noise control (target cos 0.9000) recovery 100% at τ=0.5; Gaussian-null match rate 0.0% at every τ (null max-cos p99.9 0.157–0.172 ≪ 0.5) |
| **G2 cross-seed floor** | **FAIL** — h16384 s0↔s1 firing-atom match fraction at τ=0.5 = **0.151** (bootstrap CI [0.150, 0.151]) < 0.20. Per frozen prereg: **splitting/absorption claims = INSTRUMENT_FAILURE** |
| G3 probe power (x) | **PASS** — domain 0.881 ≥ 0.50, toklen R² 0.976 ≥ 0.30, shuffled at chance (domain 0.269, family 0.197, toklen R² −0.09) |
| G4 matched budget | **PASS** — epochs=40, k=32, n_train=280000, lr=4e-4 identical incl. new h2048; x_mean/scale byte-identical (rel diff 0.0). Logged caveat: batch 2048 (not 4096) for h≥32768 per w40 VRAM rule → 2× grad steps at same data passes |

G2 cross-seed note: the smoke pass showed ~0.39–0.40 (≫ 0.20) but full = 0.151. Not a matcher bug
(mechanical controls all perfect) — on the small smoke corpus only high-frequency atoms fire, and
those ARE seed-stable; on the full 36k corpus most firing atoms are low-frequency and
seed-idiosyncratic. Lesson for future preregs: certify corpus-size-dependent gates at full n.

## Splitting vs absorption (τ=0.5, seed0, narrow atoms firing ≥20× on 042's 36k pile z) — GATED (see G2)

| pair | n_narrow | preserved | split | duplicate | partial | lost | novel (wide fire≥20) | novel (fire≥1) | novel act-mass |
|---|---|---|---|---|---|---|---|---|---|
| 2048→8192 | 2036 | 56.6% | **4.1%** | 3.9% | 11.5% | 23.9% | 75.7% | 76.8% | 53.0% |
| 8192→16384 | 7400 | 85.2% | **1.1%** | 2.1% | 3.7% | 7.9% | 51.6% | 52.7% | 39.6% |
| 16384→32768 | 12231 | 77.2% | **0.9%** | 1.8% | 3.0% | 17.1% | 53.9% | 57.6% | 42.0% |
| 32768→65536 | 15597 | 58.3% | **0.4%** | 1.4% | 2.9% | 36.9% | 61.6% | 69.9% | 53.0% |

- Splitting is essentially absent: 0/4 pairs reach the 5% bar; the split fraction *declines* with
  width (4.1%→0.4%). Narrow atoms are overwhelmingly PRESERVED (a near-duplicate direction exists in
  the wider dict) or LOST — not refined into finer partitions. Match rate at τ=0.5 (narrow→wide):
  76%/92%/83%/63%; multi-match at τ=0.5 is rare (19%/7%/6%/5%) and mostly redundancy (duplicate ≥
  split in 3/4 pairs).
- Novel wide atoms are the majority everywhere (52–77% of firing wide atoms; 40–53% of activation
  mass), robust across τ ∈ {0.3…0.7} (novel ≥ 42% even at τ=0.3).
- **BUT the G2 cross-seed floor failed**, so per the frozen prereg these two readouts are formally
  INSTRUMENT_FAILURE-gated: same-width different-seed dictionaries only match at 15–22% (h≥8192), so
  a wide atom with no narrow counterpart cannot be distinguished from ordinary seed-level dictionary
  non-identifiability. "Novelty" here is confounded with run-to-run randomness, not attributable to
  width per se. (Interesting residual asymmetry, noted not claimed: same-seed narrow→wide match at
  τ=0.5 is 63–92% while same-width cross-seed is 15–22% — same-seed width nesting is far stronger
  than seed reproducibility, but same-seed pairs share data order, so this is not a clean control.)

## Does width buy back the residual content? NO (clean, fully gated result)

FVU (seed0): cpool-val 0.446 / 0.443 / 0.449 / 0.437 / **0.458**; pile-val 0.710 / 0.722 / 0.731 /
0.727 / **0.748**; corpus-A 0.649 / 0.651 / 0.655 / 0.643 / **0.655** for h=2048/8192/16384/32768/65536.
A 32× width increase at matched budget buys **zero** FVU on every eval corpus — h2048 is already on
the plateau (confirming the early signal), and h65536 is slightly *worse*. ‖r‖/‖x‖ ≈ 0.74 at all widths.

Probe battery (corpus-A 12k, channels x̂ vs r; x-reference toklen R² 0.976, domain 0.881):

| width | x̂ toklen R² | r toklen R² | x̂ BoW AUC | r BoW AUC | x̂ domain | r domain | x̂ family | r family | x̂ voice | r voice | r beats x̂ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 2048 | 0.735 | 0.909 | 0.899 | 0.943 | 0.804 | 0.833 | 0.996 | 1.000 | 1.000 | 1.000 | 5/5 |
| 8192 | 0.713 | 0.906 | 0.887 | 0.937 | 0.786 | 0.832 | 0.995 | 1.000 | 0.999 | 1.000 | 5/5 |
| 16384 | 0.723 | 0.907 | 0.878 | 0.937 | 0.784 | 0.832 | 0.977 | 1.000 | 0.993 | 1.000 | 5/5 |
| 32768 | 0.698 | 0.910 | 0.882 | 0.930 | 0.799 | 0.841 | 0.975 | 1.000 | 0.994 | 1.000 | 5/5 |
| 65536 | 0.686 | 0.910 | 0.877 | 0.930 | 0.788 | 0.844 | 0.959 | 1.000 | 0.985 | 1.000 | 5/5 |

r beats x̂ on **5/5 certified families at every width** (042/043's finding is width-invariant). x̂'s
linear content actually *declines* mildly with width (toklen 0.735→0.686, family 0.996→0.959) while
r is flat ~0.91. Where the extra capacity goes: fire≥20 fraction collapses 99.4%→23.2% and dead
fraction rises 0%→4.0% (h2048→h65536) — width is spent on rare, seed-idiosyncratic atoms, not on
capturing the shared linearly-decodable content that lives in r.

## Cross-seed stability vs width (τ=0.5, firing atoms)

h2048 **0.473** → h8192 0.217 → h16384 0.151 → h32768 0.157 → h65536 **0.149** (s1 direction within
0.01 everywhere; median max-cos 0.450→0.200). Drop h2048→h65536 = **32.4 pts** ≥ 10. Most of the
collapse happens by h8192; beyond h16384 it is flat-low. Even at h2048, less than half the
dictionary is seed-reproducible at cos 0.5.

## Predictions (frozen) → outcomes → Brier

| Pred | P | Outcome | Brier |
|---|---|---|---|
| P1 split ≥5% in ≥3/4 pairs | 0.55 | **FALSE** (0/4; max 4.1%) | 0.3025 |
| P2 novel ≥25% every pair | 0.70 | **TRUE**¹ (min 51.6% fire≥20 / 52.7% fire≥1) | 0.0900 |
| P3 r beats x̂ ≥3/5 at h65536 AND pile FVU ≥0.60 | 0.75 | **TRUE** (5/5; FVU 0.748) | 0.0625 |
| P4 cross-seed drop ≥10 pts h2048→h65536 | 0.65 | **TRUE** (32.4 pts) | 0.1225 |

**Mean Brier (P1–P4) = 0.144.** ¹P1/P2 are the splitting/absorption claims gated by the G2
cross-seed failure — scored on raw readouts for calibration, but their *interpretation* is
INSTRUMENT_FAILURE per prereg. Ungated subset (P3, P4) mean Brier = 0.093. Internal prereg tension,
acknowledged: P4 predicted (correctly, P=0.65) a stability decline whose severe form is exactly what
tripped the ≥20% G2 floor — the gate and the prediction measured the same quantity at different
thresholds.

## Honest verdict

**Splitting does not happen at matched budget with fixed k=32** — narrow atoms carry over near-verbatim
or vanish; hierarchical refinement is <5% everywhere and shrinks with width. The complementary
"absorption" claim (novel-atom majority, 52–77%) is directionally strong but **formally
INSTRUMENT_FAILURE**: cross-seed non-identifiability (15% at h16384) means decoder-cosine novelty
cannot separate width-driven capacity from seed noise. The clean, fully-gated headline is the
**answer to the 042/043 question: width does NOT buy back the dropped content.** FVU is flat from
h2048 to h65536 on three corpora, the residual r beats x̂ on every certified probe family at every
width, and x̂'s readable content mildly degrades while dead/rare atoms proliferate. The k=32
dictionary-drops problem is a sparsity/objective property, not a capacity ceiling — more atoms at
fixed k do not fix it (row 045's k-sweep is now the sharper question).

## Limitations

- Splitting/absorption cells gated (G2 cross-seed FAIL) — see above; a follow-up should match on
  firing-pattern correlation (activation space) rather than decoder cosine to de-confound direction
  non-identifiability from functional novelty, and/or average over ≥3 seeds per width.
- Same-seed adjacent-width pairs share data order; the strong "preserved" fractions may be partly
  trajectory correlation, not pure width nesting. Not controlled here.
- Matched tokens/epochs, not FLOPs; batch 2048 for h≥32768 (2× grad steps, logged). Training corpus
  is w40 c_pool; matching/probes on SONAR pile z + corpus-A (domain shift constant across widths —
  internal comparisons only). τ=0.5 primary with {0.3–0.7} sensitivity reported; firing thresholds
  (≥20 narrow, ≥20/≥1 wide) fixed a priori. Role probe cell excluded a priori (043). Single matching
  corpus (36k); n_bow=66 words. Smoke-vs-full G2 discrepancy documented above.

## Follow-up worth funding? **Y** (modest)

Row 045 (k-scaling at fixed h) is already queued and is now the decisive arm: if FVU/residual
content is also k-invariant, the SAE objective itself caps readable content; if k buys it back, the
042/043 residual is "content beyond the top-k budget", which reframes all w40-family dictionaries.
Also cheap and worth it: one activation-correlation rerun of the novelty matcher (fixes the gated
cell with data already on the box).
