# RESULT — 066 bits-accounting: I(z; sentence) via decoder likelihoods; knee in bits

**Tier: T3-exploratory. DONE + SELF-HARVESTED in-session.** tmux c100_066 launched
17:56:10Z, DONE 17:59:52Z (3.7 min wall; forward-only — zero encoding, 063 items+z reused;
single GPU CVD=0). eng/deu/jpn × c∈{1,2,3,4,6,8} × 64 = 384 items/lang, 0 exclusions,
4 teacher-forced passes each (z_true / μ_L-ref / cross-attn-ablated prior / shuffled-z).
All 5 gates PASS (pad-consistency 5.6e-5 nats; eng c=1 I=182.4 bits ≥30; shuffled-null
0.13% of matched; KL(full‖prior)=3.98; μ/z norm ratio 0.218–0.314 logged per design).

## Verdict
★ **A SONAR vector carries ~180–195 bits about a single sentence in every language tested
— 063's huge chrF "level tax" nearly VANISHES in information units (jpn 163 vs eng 182
bits despite chrF 34.5 vs 89.7). By the preregistered measure the bits curve never
saturates (P2 FALSE) — but post-hoc subtraction of the shuffled-z generic component
reveals a clean sentence-SPECIFIC ceiling of ~460 bits (eng) / ~485 (deu) / ~390 (jpn),
reached at ~4 FLORES rows: information keeps accumulating well PAST the chrF knee
(~2 rows). String fidelity collapses before the vector stops absorbing bits.**

(a) **Bits vs length (μ-ref, preregistered measure I_mu).** Mean bits per item rise
monotonically with content and never meet the frozen saturation criterion
(m_last < 0.3·m_1 — P2 FALSE): eng 182→305→407→482→583→700 bits (c=1..8), marginals
122.8/101.7/75.2/50.5/58.7 bits/row; bits per src token fall smoothly 5.25→2.84.
Frozen-criterion knee: eng 5.0 rows / 97 src tok, deu 5.0 / 109, jpn 3.5 / 75 — LATER
than 063's chrF knees (2.14 rows / 65.3 tok eng), so P3 FALSE ([1.5,3.0] band missed).

(b) **The shuffled-z null exposes what the raw curve conflates.** Mismatched-but-real z
(roll-1 within (lang,c) = adjacent consecutive-row groups, i.e. same/nearby news articles)
contributes ≈0 at c=1 (eng 0.2 bits — the null certifies the instrument, G3) but grows to
246 (eng) / 368 (deu) / 178 (jpn) bits at c=8: at length, a large share of raw likelihood
gain is GENERIC (domain/register/length), not sentence-specific. **Post-hoc specific bits
I_mu − I_shuf saturate hard in all 3 langs:**
| rows c | 1 | 2 | 3 | 4 | 6 | 8 |
|---|---|---|---|---|---|---|
| eng | 182 | 293 | 373 | 431 | **463** | 455 |
| deu | 193 | 325 | 429 | 482 | **485** | 456 |
| jpn | 158 | 255 | 326 | 367 | **393** | 376 |
Specific marginals go ≈0/negative by c=4–6 everywhere (eng 110/81/57/16/−4). Ceiling
≈ 0.45 bits per dimension (1024-d). One FLORES row ≈ 160–195 specific bits ⇒ the ~2-row
chrF knee sits where cumulative demand (~365 bits) approaches the ~460-bit budget —
chrF collapses as headroom vanishes; I keeps rising to the cap at ~4 rows. [POST-HOC]

(c) **Cross-language budget quasi-equality (P4 TRUE, preregistered).** Ceiling I(c=8):
deu 823 / jpn 554 vs eng 700 — both within the [0.6,1.4]× band (1.18× / 0.79×); specific-
bits peaks even tighter (1.05× / 0.85×). Combined with near-equal c=1 bits, the 063 level
tax is mostly NOT missing information: jpn's z carries ~85–90% of eng's bits while
teacher-forced; free-running greedy decode + chrF's script confound is where the surface
loss happens. Information is present; verbatim decode can't surface it.

(d) **Concentration (per-token profile, full−μ gap in bits).** Front-loaded: eng
relative-position decile profile 5.10→2.32 bits/token (first/last ratio 2.2×; deu 4.01→
2.49, jpn 3.74→1.82) — consistent with 057's scaffold finding [report-only]. Word class
(eng c=1): content words 6.13 vs function words 4.03 bits/token — real but only 1.52×,
below the preregistered 2× (P5 FALSE); z's advantage over μ includes substantial
function-word/structure prediction (ties 043's function-words-in-residual result).

(e) **Reference choice audit.** Ablated-prior reference gives systematically LARGER I
(eng c=1 199.7 vs 182.4; c=8 824.8 vs 700.4) — the z-free decoder is OOD and a weak LM,
so μ-ref (identical tokenizer, in-distribution) is the conservative primary as
preregistered; conclusions unchanged under either reference.

## Predictions → Brier (frozen in PREREG_LITE)
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 instrument certifies (pos ctrl + null) | 0.90 | **TRUE** (182.4 bits; null 0.13%) | 0.0100 |
| P2 eng bits ceiling by frozen criterion | 0.70 | **FALSE** (m_last 58.7 > 36.8) | 0.4900 |
| P3 eng content knee ∈ [1.5,3.0] rows | 0.55 | **FALSE** (5.0) | 0.3025 |
| P4 deu+jpn ceiling within [0.6,1.4]× eng | 0.50 | **TRUE** (1.18× / 0.79×) | 0.2500 |
| P5 content ≥ 2× function bits/token | 0.65 | **FALSE** (1.52×) | 0.4225 |

**Mean Brier = 0.295** (worst-calibrated row so far; the P2/P3 misses share one cause —
the preregistered I_mu measure conflates sentence-specific with generic bits, and the
generic component grows with length, pushing the raw curve past every saturation test.
The post-hoc corrected curve behaves exactly as the capacity theory predicted, but that
correction was not frozen, so the misses stand as scored.)

## Limitations / what did NOT run
- I_hat is a DECODER-EXTRACTABLE lower-bound proxy for I(z;s), not true mutual
  information; all bits are "bits usable by this decoder via likelihood".
- The specific-bits ceiling (headline b) is POST-HOC. Worse, the shuffled baseline is
  roll-by-1 → adjacent article groups → I_shuf likely OVERSTATES the generic component
  (specific bits conservative). A preregistered rerun needs random cross-domain mismatch
  + item-level pairing.
- Frozen knee/saturation criteria were defined on I_mu; by construction they answered a
  different (conflated) question — scored honestly as misses.
- chrF-vs-bits comparison inherits 063's caveats (jpn chrF script confound); "info present
  but decode can't surface it" is inferred from teacher-forced vs free-running gap, not
  tested with alternative decode strategies.
- Word-class analysis eng-only, crude stopword list; per-token bits attribution ignores
  that a gain at token t may reflect information about earlier context. Positional
  profile confounds position with content (concatenated news items).
- tur/zho/arb not scored (eng primary, deu/jpn secondary per prereg); single
  encoder/decoder pair; greedy-free (likelihood-only) design.

## Follow-up worth funding? **Y (narrow).**
1. Prereg the specific-bits ladder properly: random non-adjacent cross-domain mismatched
   z, item-level pairing, all 6 langs — freeze the ceiling + knee criteria on THAT curve.
2. "Bits present, decode can't surface": beam/sampling/constrained decode on jpn near the
   knee (ties 026-decoding-strategy) — can better search recover the chrF the bits say
   is there?
3. Connect the ~460-bit ceiling to 045's k-atom curve / 028 SAE decomposition:
   bits-per-atom accounting.

## Provenance / hygiene
Repo: PREREG_LITE.md (frozen pre-compute), src/{run_066.py,run_066.sh}, out/{results_066_
full.json, scores_066_full_{eng,deu,jpn}.json, run.log}. Box dir 5 MB. GPUs verified free
pre-claim (35 MiB, hard-guard in script + shell) and free at exit (35/15/15); tmux session
self-ended, box tmux server empty; no foreign process touched. Smoke passed first
(artifacts deleted). Local commit, no push. Note for the ledger: the coordinator caught
this manager re-committing the watcher-rule slip (background waiter) mid-run — recovered
by inline polling; no data impact.
