# RESULT — 048 atom-ontogeny: when do dictionary atoms crystallize during base-autoencoder training?

**Tier: T3-exploratory.** Shape **A** ran (genuine base-model training checkpoints exist):
ladder TAE `A_P90_s0` + `A_P90_s1` (d_z=256, 12M params, 500M-token budget), 5 milestone
ckpts each at steps 0 / 2000 (5.7% of budget) / 4000 (11.4%) / 26000 (74.4%) / ~34970
(100%). Per ckpt: encode fixed corpora (320k owt_train lines SAE-train, 20k owt_val
held-out eval — identical rows for every ckpt) → train TopK SAE (h=512, k=16, e120) × 2
SAE seeds → activation-correlation matching (046's matcher, extended for
independently-trained SAEs: max Pearson corr of dense ReLU pre-act profiles over the 20k
eval rows, cross-SAE-seed primary). tmux `c100_048`, orgs parallel on CVD 0/1 (phys
0/2), launched 06:06:53Z, DONE 06:15:14Z (**8.4 min**; smoke + instrument pilot ran
first), session killed post-harvest, box intermediates deleted. All 4 gates PASS.
**Mean Brier 0.179** (P1 T, P2 F by 0.002, P3 T, P4 T).

## Verdict (headline)

**Atoms crystallize gradually and front-loaded, tracking capability, and are fully
crystallized well before training ends.** Ceiling-normalized match to the final
dictionary M(t)/C(t) at τ=0.5 (org s0): step 0 → **0.16**, step 2000 (5.7% of budget) →
**0.52**, step 4000 (11.4%) → **0.77**, step 26000 (74%) → **0.98** (at ceiling —
matching final as well as final matches itself across SAE seeds). Replicated almost
exactly in org s1 (0.14 / 0.53 / 0.76 / 1.02). No abrupt transition between milestones;
no late reorganization: the step-26000 dictionary is already indistinguishable from
final. A **frequency-first** ordering is strong at every stage, and ~16% of the
ceiling-relative inventory (~31% of the top-frequency quartile) exists **before any
training** — corpus-statistics atoms that a random encoder already supports.

## Gates

| Gate | Outcome |
|---|---|
| G1 ckpt identity | **PASS** — all 8 gated ckpt-meta val_f1 exact vs inventory; z finite; per-ckpt norms/scales logged (mean‖z‖ 5.6→8.9 untrained→trained; per-ckpt normalization applied) |
| G2 matcher cert (full n) | **PASS** — self-match 1.0; split-half assignment stability **0.9973** (n=368); scrambled-derangement null **0.000** (max corr 0.033); cross-SAE-seed ceiling C(final) **0.7227** ≥ 0.5 (median maxcorr 0.865) |
| G3 SAE quality | **PASS** — all 20 SAEs: dead 0.0%, all 512 atoms fire ≥20 on eval; FVU 0.31–0.42 (trained ckpts) |
| G4 matched budgets | **PASS** — identical cfg all 20 SAEs (h512 k16 n320k e120 lr4e-4 b1024) |

## Crystallization curve (τ=0.5 primary; cross-SAE-seed M: final-seed0 → t-seed1)

| step | %budget | val_f1 | M s0 | C s0 | **M/C s0** | M s1 | C s1 | **M/C s1** | SAE FVU s0 |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | — | 0.098 | 0.595 | **0.164** | 0.086 | 0.595 | 0.144 | 0.194 |
| 2000 | 5.7% | 0.470 | 0.295 | 0.572 | **0.515** | 0.287 | 0.543 | 0.529 | 0.310 |
| 4000 | 11.4% | 0.551 | 0.465 | 0.602 | **0.773** | 0.481 | 0.631 | 0.762 | 0.351 |
| 26000 | 74.4% | 0.704 | 0.754 | 0.766 | **0.984** | 0.752 | 0.738 | 1.019 | 0.413 |
| final | 100% | 0.709 | (ref) | 0.723 | — | (ref) | 0.748 | — | 0.416 |

- Robust across τ (0.3: 0.30→0.62→0.71→0.95 raw; 0.7 same shape) and direction
  (M_t→final within 0.05 everywhere); same-SAE-seed variant within 0.03 of cross-seed at
  every point → trajectory-correlation inflation is negligible here.
- **FVU control**: SAE FVU *worsens* monotonically as the TAE trains (0.19→0.42; z
  becomes richer/harder to compress), while match rises — crystallization is NOT
  reconstruction improving. The identifiability ceiling C(t) itself rises modestly
  (0.57→0.75 trained range): dictionaries get more seed-reproducible as features sharpen.
- Adjacent-ckpt matching shows the same shape: 2000→0 is the weak link (0.17), mid
  links ~0.50–0.51, final→26000 = 0.75 ≈ ceiling.
- Per-atom first-match (persist-to-end rule, s0, n=512): step0 9%, step2000 20%,
  step4000 18%, step26000 29%, final-only 25% — mass spread across the whole trajectory
  (gradual), no single crystallization epoch. ("final-only" includes flickerers by
  construction.)

## Frequency stratification (match to final at τ=0.5, by eval fire-count quartile, s0)

| step | q0 (rare) | q1 | q2 | q3 (frequent) |
|---|---|---|---|---|
| 0 | 0.023 | 0.016 | 0.039 | **0.313** |
| 2000 | 0.172 | 0.172 | 0.227 | **0.609** |
| 4000 | 0.422 | 0.320 | 0.367 | **0.750** |
| 26000 | 0.742 | 0.617 | 0.695 | **0.961** |

Frequent atoms crystallize first at every stage (044's matchability caveat is here a
*finding*): nearly a third of the top quartile is present in the **untrained** encoder —
these are corpus-statistics detectors that survive a random transformer projection —
and rare atoms are the last to lock in (still 26% unmatched at step 26000).

## Predictions (frozen; P1–P4 unchanged through the pre-run amendment) → Brier

| Pred | P | Outcome | Brier |
|---|---|---|---|
| P1 M(2000)/C(2000) ≥ 0.5 | 0.55 | **TRUE** (0.515; s1 0.529) | 0.2025 |
| P2 M(0) ≥ 3× null AND ≥ 0.10 raw | 0.60 | **FALSE** (0.0977 < 0.10 by 0.0023; s1 0.086) | 0.3600 |
| P3 M(t) non-decreasing | 0.75 | **TRUE** (strictly: .098/.295/.465/.754) | 0.0625 |
| P4 freq q3−q0 ≥ 10 pts at step 2000 | 0.70 | **TRUE** (43.8 pts) | 0.0900 |

**Mean Brier = 0.179.** P2 scored strictly FALSE at the frozen 0.10 threshold though the
phenomenon it predicted is clearly real (τ=0.3 raw match at step 0 = 0.30, q3 = 0.31,
scrambled null = 0.000) — a 046-P4-style near-threshold miss; direction right, number a
hair under.

## What did NOT run / deviations (all pre-logged in PREREG_LITE amendment)

- **Instrument recalibration was required**: the originally frozen SAE config
  (h2048/k32/40k/e80) had cross-SAE-seed ceiling **0.008** — dead on arrival (044's
  non-identifiability, amplified at 8× overcompleteness on d_z=256). A pre-run pilot
  (final ckpts only) selected h512/k16/320k/e120 (ceiling 0.723). All numbers above use
  the amended config; k=16 ≠ w40 reference k=32 (limitation).
- Frozen G2(i) split-half formulation was mechanically vacuous (correlated disjoint row
  sets); replaced pre-run by assignment-stability (0.997 observed).
- Time resolution is 5 points with a big 4000→26000 gap (milestones were saved at F1
  crossings, not log-spaced steps); "abrupt vs gradual" is resolved only at this
  granularity. Milestone filenames are miscalibrated (ms_f1_30 ≈ val_f1 0.47) — steps
  from ckpt meta used throughout.
- SAEs are per-checkpoint instruments on ladder-TAE z (d=256), NOT the w40 SONAR
  dictionary — extrapolation to SONAR-scale SAEs untested. Single rung (A), single arm
  (P90), 2 organism seeds. Box SAE ckpts + z deleted post-analysis (fully reproducible:
  deterministic seeds/config; results JSONs banked in `out/`).
- Matching corpus = 20k rows; certification done at full n per 044's lesson.

## Follow-up worth funding? **Y** (modest)

The 4000→26000 gap hides where the last quarter of atoms locks in; the ladder trainer
saves F1-crossing milestones, so a cheap rerun of rung A with log-spaced step ckpts
(shape B-style dense trajectory, ~70 min GPU) would resolve the tail and test the
flicker/"final-only" 25%. Also: the untrained-encoder atom inventory (~16% of ceiling,
0.31 of frequent quartile) is a clean handle on "what SAEs find that isn't learned" —
worth cross-referencing with 047's topic-selectivity labels.
