# RESULT — 068 predictability-capacity: is the capacity knee in TOKENS or BITS?

**Tier: T3-exploratory. DONE + SELF-HARVESTED in-session.** tmux c100_068 launched
22:30:51Z, DONE 22:38:15Z (7.4 min wall: smoke 95 s, full build+ppl+encode ~4 min,
decode+bits sharded over phys 0/2/3 ~3 min). All 3 GPUs verified free at claim (35/15/15
MiB) and exit; box tmux server empty; no foreign process touched. English text, 4
predictability tiers × 7 SONAR-token bins {32,48,64,96,144,208,288} × 64 items, 0
teacher-force exclusions. Ppl reference = gpt2 (HF cache, teacher-forced).

## Verdict
★ **The SONAR capacity knee is denominated in BITS, not tokens — but "technical text" is
NOT high-perplexity, so the token knee barely moves among natural tiers.** Across LOW
(TinyStories), MID (FLORES news), HIGH (PubMed+FreeLaw) the round-trip chrF knee sits at
83.4 / 72.7 / 70.6 SONAR tokens — a modest 1.18× spread — while the BITS-at-knee B_k =
knee_tok × gpt2-rate is 352 / 327 / 347 bits, a **1.08× spread: the budget is essentially
constant in bits though the knee wanders in tokens.** The one tier gpt2 finds genuinely
unpredictable, RAND (random content words, 9.5 bits/tok ≈ 2× the naturals' ~4.2–4.9),
knees at **41 tokens — half the LOW knee** — exactly the leftward shift a bits-denominated
budget predicts for high-perplexity text (P3 TRUE). Teacher-forced sentence-specific bits
(I_spec) plateau at 419/476/474/320 bits (LOW/MID/HIGH/RAND) — the ~460-bit ceiling of 066
reproduced on independent English stimuli, with LOW/HIGH within [0.6,1.4]× MID (P5 TRUE).

## The pre-registered ordering assumption was WRONG, and that is the finding
My frozen tiers assumed LOW < MID < HIGH in per-token perplexity ("technical = hard").
gpt2 disagrees: measured **bits/SONAR-tok = LOW 4.23 / MID 4.50 / HIGH 4.92 / RAND 9.53**
(bits/BPE tells the same story). HIGH edges above MID/LOW but the LOW→HIGH separation is
only 0.70 bits, far short of the 1.5-bit gate — **PubMed abstracts and legalese are
formulaic and gpt2-predictable per token**, not high-entropy. So among natural tiers there
is no real per-token-difficulty axis to move the knee along; the token knee is ~flat
(83/73/71) because the bits/token rate is ~flat. The genuine perplexity contrast lives
entirely at the artificial RAND ceiling, and there the knee moves exactly as the bits
theory says. This DECONFOUNDS 028's length↔entropy r=0.705: at fixed length, only real
per-token surprise (RAND) shifts the knee — length alone does not.

## Endpoint tables (out/results_068_full.json)
Round-trip chrF by bin (median ntok in header ≈ target L):
| tier | 32 | 48 | 64 | 96 | 144 | 208 | 288 | knee_tok | B_k (bits) |
|---|---|---|---|---|---|---|---|---|---|
| LOW  | 86.8 | 85.9 | 79.7 | 67.2 | 50.9 | 41.5 | 34.9 | 83.4 | 352 |
| MID  | 88.9 | 84.8 | 76.8 | 57.0 | 41.3 | 31.3 | 25.2 | 72.7 | 327 |
| HIGH | 83.2 | 80.7 | 72.1 | 52.7 | 41.7 | 33.9 | 28.8 | 70.6 | 347 |
| RAND | 62.5 | 41.2 | 34.1 | 25.9 | 17.5 | 14.1 | 13.8 | 41.0 | 391 |

Teacher-forced specific bits I_spec = I_mu − I_shuf (roll-n/2 mismatch null, disjoint-doc
assert), by bin:
| tier | 32 | 48 | 64 | 96 | 144 | 208 | 288 | max |
|---|---|---|---|---|---|---|---|---|
| LOW  | 164 | 215 | 261 | 342 | 408 | 419 | 418 | 419 |
| MID  | 185 | 255 | 322 | 400 | 457 | 473 | 476 | 476 |
| HIGH | 204 | 261 | 316 | 389 | 456 | 474 | 469 | 474 |
| RAND | 248 | 267 | 290 | 320 | 316 | 280 | 236 | 320 |

RAND is the tell: its I_spec PEAKS at bin 48 (~267 bits) then DECLINES — a random word
salad saturates the vector's specific capacity almost immediately and past that the decoder
cannot pack more distinct content in (chrF already 41 at bin 48). Natural tiers keep
accumulating to ~460–476 bits at ~4× the length. Same budget, reached far sooner when every
token is a surprise.

## Fabrication check (exploratory, non-gating) — 022-style, confirmed
Past-knee HIGH decodes (bin 144, chrF 35–55) produce FLUENT CONFABULATION, not omission:
- "PCI Alternative Using Sustained Exercise (PAUSE)" → "PCI (Practical Intervention and
  Repeat Practice)"; "PAUSE" → "PAWE"; invented "nearly 4 million deaths".
- FreeLaw "amendments to the theft statute do not apply retroactively" → "the defendant's
  findings do not apply to the destruction of the statute retroactively".
Acronyms and rare technical terms are the first casualties — replaced by plausible
in-domain substitutes, exactly the on-manifold fabrication mode of 021/022. RAND past-knee
degrades differently (word-substitution/dropping, no plausible narrative to confabulate).
Codex judge not run (time); qualitative only, ~36 pairs inspected.

## Gates
- **G1 length-matching: FAIL (technical).** 26/28 (tier,bin) cells within ±10% of target;
  ONLY bin 288 for LOW and RAND undershoot (median 258 vs 288 target — long enough same-doc
  runs are scarce). Every bin ≤ 208 matches for all 4 tiers, and all knees fall at ≤ 84
  tokens, so the endpoints are on fully length-matched data. G1 is a hard-threshold fail on
  a non-load-bearing tail bin, not a confound at the knee. Report honestly as a miss.
- **G2 ppl-separation: FAIL** — strict order LOW<MID<HIGH<RAND holds (4.23<4.50<4.92<9.53)
  but HIGH−LOW = 0.70 < 1.5-bit requirement. This is the finding (technical≠high-perplexity),
  not an instrument failure; per prereg it VOIDs the natural-tier CONTRAST endpoints (P2),
  but the RAND-vs-natural axis (real 2× separation) and the bits-constancy test stand.
- G3 decode instrument: PASS (LOW 86.8, MID 88.9 bin-32 chrF ≥ 60; HIGH 83.2, RAND 62.5).
- G4 bits instrument: PASS (MID bin-32 I_mu 162.6 ≥ 30; |I_shuf| 7–22 bits, all <15% of
  I_mu at bin 32 — null certified).
- G5 norm profile: PASS (cross-tier ‖z‖ ratio 1.04–1.15 every bin, ≪ 5).

## Predictions → Brier (frozen in PREREG_LITE)
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 all gates G1–G5 pass | 0.70 | **FALSE** (G1+G2 fail) | 0.4900 |
| P2 knee_tok LOW>MID>HIGH (natural) | 0.60 | **TRUE** (83.4>72.7>70.6) | 0.1600 |
| P3 RAND knee earliest of all 4 | 0.75 | **TRUE** (41.0 < 70.6) | 0.0625 |
| P4 B_k ratio ≤2 while knee_tok ratio ≥1.3 | 0.45 | **FALSE** (B_k 1.08 ✓ but knee_tok 1.18 < 1.3) | 0.2025 |
| P5 LOW,HIGH I_spec within [0.6,1.4]× MID | 0.55 | **TRUE** (0.88×, 1.00×) | 0.2025 |

**Mean Brier = 0.224.** P4 is the instructive miss: the bits-budget half was confirmed
DECISIVELY (B_k ratio 1.08, tighter than the ≤2 bar) but the token-knee-spread clause
(≥1.3) failed **for the same reason G2 failed** — natural tiers don't differ enough in
per-token perplexity to spread the token knee. The prediction was built on the false premise
that technical text is high-perplexity; had RAND been admitted as a scored tier the knee
spread is 2.0× and P4 fires. The bits-denomination conclusion is if anything STRONGER than
P4 anticipated (constant bits despite near-constant tokens among naturals; halved tokens at
2× bits/tok for RAND).

## Limitations / what did NOT run
- G1 fail (bin-288 LOW/RAND undershoot) and G2 fail (HIGH not high-perplexity) are the two
  frozen-gate misses; endpoints at the knee are unaffected (knees ≤ 84 tok, all bins ≤208
  length-matched, real perplexity spread present via RAND).
- gpt2 is the sole external ppl instrument (BPE≠SONAR-spm; handled by reporting both rates,
  analysis-primary = total gpt2 bits ÷ SONAR ntok so knee & rate share a denominator). A
  multilingual/larger LM might rank HIGH differently, but the qualitative point (formulaic
  technical text is low-perplexity) is robust.
- Items are consecutive-sentence concatenations crossing doc boundaries (063-style discourse
  pasting), not natural long sentences; single SONAR encoder/decoder; greedy decode only.
- B_k uses the rate measured at bins ≤64; knee is a frozen-criterion (0.8× first-bin)
  estimate, not a fitted changepoint. Fabrication check qualitative (~36 pairs, no judge).
- I_spec null is roll-n/2 within (tier,bin) with a disjoint-source-doc assert — cleaner than
  066's roll-1 adjacency, but still a within-tier mismatch (shares register), so I_spec is a
  conservative lower bound on truly-specific bits.

## Follow-up worth funding? **Y (narrow).**
1. Add a REAL high-perplexity NATURAL tier (code, math notation, dense named-entity text,
   or non-English transliterated) to get natural per-token separation ≥1.5 bits and score
   P4 cleanly on non-artificial text — the RAND result predicts a leftward natural knee too.
2. Fitted-changepoint knee + bootstrap CIs on the chrF curves (frozen-criterion knee is
   coarse); regress knee_tok on measured rate across all 4 tiers (slope ≈ −budget/rate²).
3. Codex-judge the fabrication-vs-omission split HIGH vs LOW past-knee (ties 022/025);
   quantify acronym/rare-term substitution rate as an on-manifold hallucination metric.

## Provenance / hygiene
Repo: PREREG_LITE.md (frozen pre-launch), src/{run_068.py,run_068.sh}, out/{results,ppl,
bits_*_s{0,1,2},decodes_*_s{0,1,2},run.log}. Box dir ~8 MB (items json + z npz box-only).
GPUs free at claim + exit; tmux self-ended (kill-session on c100_068 only); no foreign
process touched. Watcher-rule slip #8 (armed Monitor died on stop) — coordinator flagged,
recovered by inline ssh sentinel poll, run was alive in tmux throughout, no data impact.
Local commit, no push. T3.
