# RESULT — 067 bottleneck-dim-ladder: knee & bits vs d_z (16/64/256/1024)

**Tier: T3-exploratory. DONE.** tmux `c100_067` launched 20:29:36Z, DONE 21:44:45Z
(~75 min wall). 3 NEW full-budget trains d_z ∈ {16, 64, 1024} (ladder rung-A arm-D,
identical args to A_D_s0 except --d_z; one per A4000; all 3 rc=0, pct_budget 0.9637,
step 37360 ≡ A_D_s0's 37360/500M tok — matched budget confirmed; no DIVERGED → **G2
PASS**) + reused A_D_s0 (d=256 primary) + A_D_s1 (d=256 seed anchor). Eval rc=0 on all
5 models: 192 owt_val chunks (offset 2048+, disjoint from trainer val), nested prefixes
L ∈ {8,12,16,24,32,48,62}, greedy chrF knee (τ = 0.8·chrF(L=8), 063-frozen) + teacher-
forced I_mu / I_spec with roll-n/2 shuffled null (066 amendment). **G1 PASS** (GPUs
35/15/15 MiB at claim, in-script guard). **G4 PASS**: lp_true > lp_shuf on 100% of items
at every bin for ALL five models (gate required ≥95% at L=8 for d ≥ 256).

## Headline curves (all knees uncensored)

| model | d_z | L=8 chrF | L=62 chrF | knee (tok) | I_spec(L=62) bits | bits/dim | PR | val_f1 |
|---|---|---|---|---|---|---|---|---|
| dz16 | 16 | 28.44 | 20.09 | 20.45* | 109.6 | 6.85 | 14.8 | 0.233 |
| dz64 | 64 | 81.79 | 27.12 | **13.42** | 319.2 | 4.99 | 60.3 | 0.496 |
| dz256_s0 | 256 | 84.52 | 34.27 | **20.97** | 404.0 | 1.58 | 219.2 | 0.655 |
| dz256_s1 | 256 | 84.50 | 34.96 | 21.34 | 409.2 | 1.60 | 217.7 | 0.649 |
| dz1024 | 1024 | 85.08 | 36.24 | **21.32** | 404.2 | 0.39† | 655.3 | 0.666 |

\* dz16's relative-τ knee is DEGENERATE — see verdict. † rank ceiling 256: per
*effective* dim dz1024 is 404/256 = 1.58, identical to dz256.

## Verdict

★ **By the frozen criterion the verdict is OBJECTIVE-limited (dim\* = 16 ≤ 64) — but
that branch fires only through a degenerate statistic, and the weight of evidence says
the opposite: absolute fidelity and specific bits are CAPACITY-limited from 16→64→256
and saturate exactly at the architectural rank ceiling, with the d_z=1024
reparameterization placebo landing dead flat (knee 21.32 vs 20.97, excess 0.35 ≤ seed
gap 0.37).**

- **Frozen letter.** max knee = 21.34, 0.9·max = 19.21; smallest d_z with knee ≥ 19.21
  is 16 (20.45) → dim* = 16 → "OBJECTIVE-limited". Recorded as scored. **Instrument
  caveat (flagged, not overridden):** the relative-τ knee assumes a healthy L=8 anchor.
  dz16 fails the P1-style control (L=8 chrF 28.4 « 60, near-floor flat curve 28→20), so
  its τ = 22.8 is crossed late for the wrong reason. Knee is non-monotone in d_z
  (dz64 = 13.4 < dz16 = 20.5), so "smallest d_z within 0.9·max" is ill-defined in
  spirit. The criterion, as frozen, did not anticipate a rung failing its own anchor.
- **[POST-HOC] absolute-τ knee (chrF = 50 crossing, same data):** dz16 never reaches 50;
  dz64 18.9; dz256 32.5/32.0 (s0/s1); dz1024 36.3. Monotone in d_z, ratio 256/64 = 1.71,
  1024/256 = 1.12 (≤ 1.25 placebo band) → capacity-limited up to the rank ceiling, then
  flat. Not frozen, so it does not overturn the scored verdict; it explains it.
- **Bits ladder (secondary, frozen).** I_spec(L=62): 109.6 → 319.2 → 404.0 → 404.2
  (16→64→256→1024). log2-log2 slope over 16→256 = **0.47** — between the frozen
  objective-ish (≤0.2) and capacity-ish (≥0.6) poles, leaning capacity; scaling is
  strongly sublinear (bits/dim FALLS 6.9 → 5.0 → 1.6). 256→1024 gains +0.2 bits
  (+0.06%): the placebo is exactly flat in specific bits, as the rank argument demands.
- **Seed anchor.** s0 vs s1 (d=256): knee 20.97 vs 21.34 (Δ0.37), L=8 chrF Δ0.02,
  I_spec(62) Δ5.2 bits (1.3%) — seed noise is small; all cross-dim gaps above except
  the placebo's are many× the seed gap.
- **Placebo fine print.** dz1024 shows a small optimization benefit beyond seed range in
  val_f1 (0.666 vs 0.655/0.649) and L=62 chrF (+1.3 over s1) — consistent with an easier
  optimization landscape for the overparameterized to_z head, NOT extra capacity
  (I_spec flat); well inside the ≤1.25× knee band, so no MIXED/optimization flag.

**Connection to 066 (qualitative only — different model family: this is the ladder
BART-DAE TAE on owt token-prefixes, not SONAR on FLORES).** 066 found SONAR saturating
at ~460 specific bits ≈ 0.45 bits/dim at d=1024. Here the d_z=1024 rung shows 404 bits
≈ 0.39 bits/dim — numerically close but coincidental: 067's encoder is rank-256, so the
honest per-effective-dim figure is 1.58 bits/dim, and the sublinear slope (0.47) shows
bits/dim is not a family constant but falls with width. The shared qualitative story
holds in both families: specific bits saturate at a model-set ceiling while the raw
I_mu curve keeps growing (dz16: I_mu 266.6 vs I_spec 109.6 at L=62 — the roll-n/2 null
absorbs a large generic component, 066's lesson applied as amended).

## Predictions → Brier (frozen in PREREG_LITE)

| pred | P | outcome | Brier |
|---|---|---|---|
| P1 d=256 pos ctrl: L=8 chrF ≥ 60, decreasing, knee uncensored | 0.85 | **TRUE** (84.5; 84.5→34.3 monotone; knee 21.0 < 62) | 0.0225 |
| P2 knee(256_s0)/knee(16) ≥ 2.0 | 0.75 | **FALSE** (20.97/20.45 = 1.03) | 0.5625 |
| P3 placebo knee(1024)/knee(256_s0) ≤ 1.25 | 0.65 | **TRUE** (1.017; excess 0.35 ≤ seed gap 0.37) | 0.1225 |
| P4 slope log2 I_spec(62) vs log2 d_z ≥ 0.5 over 16→256 | 0.60 | **FALSE** (0.47) | 0.3600 |
| P5 L=8 chrF non-decreasing in d_z (−2 tol) | 0.80 | **TRUE** (28.4→81.8→84.5→85.1) | 0.0400 |

**Mean Brier = 0.2215.** G3 (=P1) PASS — instrument certified. The P2 miss is the same
event as the degenerate verdict: the relative-τ knee was expected to shrink at low d_z,
but a collapsed anchor shrinks τ in proportion, cancelling the effect. P4 missed by
0.03 of slope — real sublinearity, prediction was directionally right.

## Limitations / what did NOT run
- **Rank ceiling above 256**: encode_tokens mean-pools to d_model=256 before to_z, so
  capacity vs objective is architecturally indistinguishable above d_z=256 (prereg-
  stated); the 1024 rung tests only reparameterization, and does so cleanly.
- **Relative-τ knee invalid at d_z=16** (anchor collapse) — the frozen verdict branch
  "dim* ≤ 64" is triggered by this artifact; absolute-τ reading is post-hoc.
- 1 seed per new dim (seed noise gauged only at d=256); max_len=64 right-censors knees
  ≥ ~62 tok (none censored in practice); ladder-family caveat — this is the ladder TAE,
  NOT SONAR; the 066 comparison is qualitative. Full budget ran (no short-budget caveat).
- I_spec remains a decoder-extractable lower-bound proxy, not true MI.

**Follow-up worth funding? Y (cheap).** (a) Freeze the absolute-τ knee (or a bits-demand
knee) as the primary and add d_z ∈ {32, 128} rungs (~3.4 GPU-h) to localize the
saturation point of absolute fidelity between 64 and 256; (b) a d_model-widened control
(rank ceiling moved to 512) would separate objective from architecture above 256 — the
one question this design provably cannot answer.

Artifacts: `out/results_067.json`, `out/run_067.log` (box:
`campaign100/067-bottleneck-dim-ladder/{runs/dz{16,64,1024}_s0/final.pt, out/, run.log,
DONE}`; runs/ hold final ckpts + config/metrics/train.log only — milestone cleanup
verified, ~91 MB total).
