# RESULT — 055 coconut-latents (EXT: Coconut continuous-thought vectors through the battery)

**Tier: T3-exploratory. DONE-partial** (primary GSM8k arm complete + scored; secondary ProsQA arm
killed by its own pre-registered manipulation gate — see below). Harvested 2026-08-03 by a
repair+harvest agent after the box run left out/FAILED. GPU phys0 (CVD=0) only, hard-guarded;
tmux c100_055 self-ended; wall ~14 min (11:15:34 → 11:29 UTC). Checkpoint provenance: official
facebookresearch/coconut releases NO weights (issue #37 open) — ran on the best-provenance
community/paper reproductions, Dilgren & Wiegreffe 2026 (arXiv 2604.04902, MIT; official coconut
framework @ pinned commit 27273cb): `connordilgren/gpt2-gsm8k-coconut` checkpoint_33 (primary)
and `connordilgren/gpt2-prosqa-coconut` checkpoint_40 (secondary). Pre-declared deviation (054
Mimir pattern): claims attach to "GPT-2 Coconut trained with the official code", not the paper
authors' exact checkpoints.

## Verdict (one line)
**Latent chain-of-thought pressure does NOT install binding — closing the 051–055 "does any
pressure install binding?" arc as a clean NO**: on a functionally-certified GSM8k Coconut
(0.35 greedy accuracy, paper-level), the primary binding cell (agent_patient cross-construction
+ lexical holdout) is dead chance at every latent step (0.494–0.515, max CI-lo 0.502, vs the 0.6
line) and indistinguishable from both anchors (emb_mean 0.495, base_h 0.515); within-ceiling
mildly DECAYS across thoughts (ap 0.648 → 0.602 mlp) instead of enriching; the ProsQA arm was
refused certification by the frozen manipulation gate (latent thoughts collapse to a runaway
near-fixed direction on English stimuli: consecutive-step cos 0.993→1.000, norm 182→281,
pc1_evr 0.76) so P5 is UNSCOREABLE. The arc's highest binding allowance (P(binding)=.30) did
not pay.

## What actually ran / what failed
- Phase 0–1 (night8 venv, GPU0): ckpts load exactly in Coconut save format (vocab 50260, 0
  missing keys). Functional gate gsm: 21/60 = **0.350** greedy ≥ 0.15 (paper ~34%) PASS.
  Extraction: 3200 battery sentences + 196 bag words as official eval prompts
  (`"<sentence> Who did what?\n"` + `<|start-latent|>` + 6×`<|latent|>`, manual latent feedback
  == coconut.py inference). Reps (d=768): emb_mean, base_h (unmodified gpt2 anchor),
  thought1..6. Manipulation gate gsm PASS (consecutive-thought cos 0.713→0.978, item_std
  0.26–0.49).
- Phase 2: stimuli_v2 battery UNCHANGED (051 machinery; agent_patient + genitive, linear+mlp,
  seeds 0,1,2, n_boot 1000) × 8 gsm reps — **8/8 completed, 0 errors**, sentinel DONE_GSM.
- Phase 3 (**the FAILURE**): prosqa functional gate PASSED strongly (98/100 = 0.980 ≥ 0.60),
  extraction ran, but the frozen manipulation gate (ii) FAILED: consecutive-thought cos
  ['0.993', '0.999', '1.000', '1.000', '1.000'] vs the `all < 0.999` criterion →
  `extract_states_055.py` exits 4 by design → run.sh `fail "extract prosqa"` → out/FAILED.
  **Not a bug — a legitimate gate trip**; relaxing it post-hoc would unfreeze the prereg, so the
  7 prosqa batteries were NOT run and no rerun was attempted. The prosqa states + geometry are
  on the box for any future re-design.
- Artifacts in repo out/: 8× binding_battery_coconut_gsm_*.json, BATTERY_RESULTS_055.md,
  func_check_{gsm,prosqa}.json, geometry_055_{gsm,prosqa}.json, run.log. States (.npz, ~138 MB)
  kept on box only.

## Headline numbers — GSM arm (ap = agent_patient; within = linear/mlp ceiling; primary =
cross-construction + lexical holdout, best readout; strict 0.9 gate = INSTRUMENT_FAILURE for
ALL 8 reps — genitive within max 0.698 < 0.9 — per prereg (iii)/052 A1 the science is the
deltas vs the measured emb_mean/base_h anchors)

| rep | ap within | ap primary [CI] | ap surf-cross | gen within | gen primary [CI] | z_bag |
|---|---|---|---|---|---|---|
| emb_mean (anchor) | 0.609/0.638 | 0.495 [0.487,0.502] | 0.604 | 0.500/0.500 | 0.500 [0.495,0.503] | 0.500 |
| base_h (anchor) | 0.564/0.647 | 0.515 [0.502,0.529] | 0.579 | 0.545/0.670 | 0.487 [0.460,0.512] | 0.500 |
| thought1 | 0.608/**0.648** | 0.504 [0.494,0.515] | 0.591 | 0.600/0.698 | 0.517 [0.493,0.541] | 0.500 |
| thought2 | 0.577/0.609 | 0.502 [0.488,0.515] | 0.581 | 0.580/0.650 | 0.488 [0.467,0.507] | 0.500 |
| thought3 | 0.581/0.620 | 0.506 [0.494,0.519] | 0.581 | 0.566/0.657 | 0.512 [0.485,0.544] | 0.500 |
| thought4 | 0.569/0.614 | 0.502 [0.487,0.518] | 0.575 | 0.576/0.659 | 0.510 [0.479,0.539] | 0.500 |
| thought5 | 0.568/0.598 | 0.497 [0.483,0.510] | 0.576 | 0.572/0.644 | 0.517 [0.486,0.541] | 0.500 |
| thought6 | 0.566/0.602 | 0.497 [0.483,0.510] | 0.577 | 0.568/0.637 | 0.515 [0.494,0.537] | 0.500 |

- **No binding at any latent step**: every ap-primary cell sits at chance with CI-lo ≤ 0.502
  and CI-hi ≤ 0.529 — nowhere near the 0.6 line, and flat vs both anchors (thought deltas vs
  emb_mean: −0.001…+0.011). Reasoning-through-latents adds NOTHING transferable.
- **Trajectory decays, doesn't enrich**: ap within best-readout 0.648 (t1) → 0.602 (t6), i.e.
  −0.046; linear −0.042. Later thoughts drift toward task-specific content, away from the
  sentence (direction as predicted by P3, magnitude 0.004 under the frozen −0.05 line).
- Exploratory: emb_mean's genitive positive control is EXACTLY 0.500 (a bag of gpt2 input
  embeddings carries zero within-construction genitive role info) while thought1 reaches
  0.600/0.698 and base_h 0.545/0.670 — contextual processing (trained or not) adds
  within-construction separability, but none of it survives the transfer cells. Surface-cross
  stays 0.58–0.61 everywhere (the usual order code), flipped-parity CIs straddle 0.5.

## Geometry / manipulation (geometry_055_{gsm,prosqa}.json)
- **P4 scale mismatch confirmed**: gsm thought-norm 45.0 (t1) – 56.6 (t5) vs mean-pooled input
  embedding norm 1.77 → ratio **25.4×** (≥ 5). Coconut really does feed back hidden states an
  order of magnitude larger than the embeddings GPT-2 was trained on; gsm thoughts stay
  moderately structured (anisotropy meancos 0.90–0.96, pc1_evr 0.26–0.35; base_h pc1_evr 0.99).
- **ProsQA latent collapse (the failure's cause, and a finding)**: on English battery stimuli
  the prosqa ckpt's thoughts converge to a shared runaway direction — norm grows monotonically
  182 → 281 (t1→t6), anisotropy meancos 0.9993 → 0.9999, pc1_evr 0.56 → 0.76, consecutive-step
  cos hits 1.000 by t3. The 6-stage/c_thought=1 ProsQA model (trained on synthetic ontology
  graphs only) treats OOD prose as a fixed-point attractor; "thought_t" are not distinct
  manipulations there, so battery cells would have been uninterpretable — the gate did its job.

## Gate outcomes
(i) functional: gsm 0.350 ≥ 0.15 PASS; prosqa 0.980 ≥ 0.60 PASS. (ii) manipulation: gsm PASS;
prosqa **FAIL** (→ arm not certified, batteries not run). (iii) strict 0.9 battery gate:
INSTRUMENT_FAILURE ×8 (expected; measured-anchor logic used). (iv) z_bag role-AUC = 0.500
everywhere PASS.

## Brier vs frozen PREREG_LITE.md predictions
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 thought1 ap within ≥ 0.70 | 0.70 | NO (0.648 mlp) | 0.490 |
| P2 no step primary CI-lo > 0.6 (gsm) | 0.70 | YES (max CI-lo 0.502) | 0.090 |
| P3 within t6 ≤ t1 − 0.05 | 0.55 | NO (−0.046, razor-thin miss) | 0.303 |
| P4 norm ratio ≥ 5 | 0.70 | YES (25.4×) | 0.090 |
| P5 prosqa matches P2 | 0.65 | **UNSCOREABLE** (manipulation gate failed; batteries never ran) | — |
**Mean Brier = 0.243 over the 4 scoreable predictions.** Miss pattern: over-confident on the
ceiling (P1 — GPT-2-small reps probe lower than SONAR-class codes; ceilings 0.60–0.65 mirror
051's small-embedder band), P3 missed by 0.004.

## Limitations / what did NOT run
- **Community-reproduction provenance**: checkpoints are Dilgren & Wiegreffe's official-code
  retrainings, not Hao et al.'s (withheld); claims = "GPT-2 Coconut trained with official code
  on GSM8k-Aug / ProsQA".
- **ProsQA arm did NOT run** (7 batteries): manipulation gate refused certification (latent
  collapse on OOD stimuli, above). P5 UNSCOREABLE — this is a scope reduction, not evidence
  about prosqa binding. A prosqa-native stimuli redesign could un-gate it but was out of prereg.
- Battery sentences are OOD for both models (math-word-problem / ontology training dists);
  functional certification is on the training task, not the stimuli. Within-ceilings 0.60–0.70
  → strict INSTRUMENT_FAILURE, anchors carry the inference. Single checkpoint per arm; GPT-2
  scale only (llama-3.2-1b coconut repros exist, untested).
- Follow-up worth funding? **N as a row** — the 051–055 arc is closed (five pressures, five
  nulls on transfer). Cheap fold-ins only if the theme resurfaces: (a) llama-3.2-1b-gsm8k
  coconut (scale check, same script), (b) prosqa-formatted stimuli to un-gate the collapsed
  arm.

## Arc closer — "does any pressure install binding?" (051–055)
051 embedder sweep (capacity/objective: NO) → 052 LASER/LaBSE (massively-multilingual
translation objective: NO) → 053 instruction embedders (inference-time conditioning: NO) →
054 LCM/Mimir (concept-level planning: NO) → **055 Coconut (end-to-end latent chain-of-thought,
the strongest candidate, P(binding)=.30 allowance: NO)**. Across every pressure tested, pooled/
state codes carry surface order + within-construction separability, never a transferable role
variable. Binding remains untraced to any training or inference pressure in this program.

## Ops footprint
Box dir 145 MB (states npz; no venv of ours — night8 venv reused read-only); HF cache +~1.7 GB
(2 ckpts + base gpt2, shared). Disk 553 GB free. tmux c100_055 self-ended before harvest; no
foreign process/session touched; no rerun launched (gate trip ≠ bug). Repo carries
PREREG_LITE.md, src/ (box copy authoritative), out/ JSONs + MDs + run.log.
