# RESULT — 054 lcm-audit (EXT: LCM intermediate concept states through the battery)

**Tier: T3-exploratory.** Self-harvested in-session 2026-08-03. GPU phys0 (CVD=0) only,
hard-guarded; tmux c100_054; smoke (extraction + 6-item roleswap + 2-rep battery) passed before
full run; full run wall ~9 min, **0 reps errored**, sentinel DONE. Feasibility phase found Meta's
official LCM weights UNOBTAINABLE (repo issue #4: "no plan to open source the weights") — ran on
**Mimir-1.6B** (`mimir-lcm/Mimir-1.6B`, MIT, arXiv 2605.25263), a reproduction of the official
`two_tower_diffusion_lcm_1_6B` arch (5-layer causal context tower d=2048 + 13-layer denoiser,
SONAR concept space) trained with the official facebookresearch/large_concept_model codebase on
fineweb-edu-350BT + fineweb-2. Pre-declared deviation: claims are about "a 1.6B two-tower
diffusion LCM (Mimir repro)", not Meta's exact checkpoint.

## Verdict (one line)
**The LCM context tower neither enriches nor destroys SONAR's role code — all three internal
depths sit exactly in the SONAR regime (ap within 0.74–0.80, primary binding cell dead chance at
every depth; P3 confirmed) with the transferable surface-order code mildly AMPLIFIED with depth
(surf-cross 0.755 → 0.791); next-concept prediction preserves roles at the TEXT level mostly by
echoing the role sentence (judged consistency 92% upper-bound, active role inversions 7.8%),
while the pre-registered embedding forced-choice tracking metric came back chance (0.465) and is
INVALIDATED as a role readout by a surface-subject confound (it "tracks" only 23% on echoes that
textually preserve roles).**

## What actually ran
- Stage 1 (venv_lcm: py3.11, torch 2.5.1+cu121, fairseq2 0.3.0rc1, sonar-space 0.3.2 — separate
  venv, night8 untouched): SONAR-encoded 3200 battery sentences + 196 bag-words; each pushed
  through Mimir's context tower as a 1-sentence context; states dumped at `prenet`
  (frontend out), `ctx_mid` (layer 3/5), `ctx_final` (tower output, what the denoiser
  conditions on) + `z_raw` baseline. 31 s.
- Stage 2: role-swap causal cell — 100 two-sentence contexts (S1 mention-order counterbalanced,
  S2 = "The X <verb> the Y." vs swap), ONE next concept each (seed 42, Mimir README options,
  max_gen_len=1), SONAR-decoded. 200 generations, 26 s.
- Stage 3 (night8 venv): stimuli_v2 battery UNCHANGED (051 machinery) on the 4 reps;
  agent_patient + genitive, linear+mlp, seeds 0,1,2, n_boot 1000.
- Harvest: codex judge (LOCAL, neutral framing, stdin closed, 8-way, 277 s) on all 200 records.
- Artifacts: `out/` binding_battery_lcm_{z_raw,prenet,ctx_mid,ctx_final}.json,
  BATTERY_RESULTS_054.md, roleswap_054.json, judge_054.json, geometry_054.json, run.log.

## Headline numbers — battery (ap = agent_patient; within = lin/mlp; primary = cross-constr +
lexical holdout, best readout; strict verdict INSTRUMENT_FAILURE for ALL reps incl. z_raw — same
as SONAR itself on stimuli_v2, per 052 A1 the science is the gate-relative deltas)

| rep | ap within | ap primary [CI] | ap surf-cross | ap surf-flip | gen within | gen primary | z_bag |
|---|---|---|---|---|---|---|---|
| z_raw (anchor) | 0.758/0.796 | 0.508 [0.489,0.527] | 0.755 | 0.734 | 0.827/0.835 | 0.500 [0.472,0.531] | 0.500 |
| prenet | 0.736/0.795 | 0.510 [0.492,0.527] | 0.737 | 0.713 | 0.817/0.832 | 0.501 [0.472,0.532] | 0.500 |
| ctx_mid | 0.774/0.799 | 0.501 [0.478,0.524] | 0.788 | 0.765 | 0.805/0.835 | 0.522 [0.492,0.556] | 0.500 |
| ctx_final | 0.769/0.801 | 0.501 [0.480,0.523] | **0.791** | 0.770 | 0.792/0.828 | 0.533 [0.505,0.560] | 0.500 |

- **Anchor replication is exact**: z_raw ap within 0.758/0.796 = SONAR's 052-table values to the
  third decimal. Anchor gate (≥0.70) PASSED; instrument certified for comparisons.
- **No ceiling shift**: max within delta vs z_raw = +0.016 lin / +0.005 mlp (< the +0.03 line).
  The context tower is ~role-information-preserving at every depth; multi-sentence "planning
  pressure" does not enrich (or erase) the per-sentence role code.
- **No binding emerges inside the LCM**: ap primary chance at all depths (all CI-lo < 0.53).
- **Surface code amplified with depth** (exploratory): ap surface-cross 0.755 → 0.791, and the
  parity-flipped surface cell rises in step (0.734 → 0.770) — the tower makes the transferable
  ORDER code slightly more linearly available while abstract role stays absent.
- Exploratory, not significant vs multiplicity: gen primary creeps to 0.533 [0.505,0.560] at
  ctx_final (CI-lo barely >0.5); flag-only, would need a dedicated rerun to claim.

## Role-swap causal cell (100 contexts × 2 variants)
- **Manipulation check PASS**: 98% of decoded prediction pairs differ; mean cos(pred_v1,pred_v2)
  = 0.797 — generation is context-sensitive.
- **Pre-registered forced choice: 0.465 (93/200), binom p(one-sided)=0.856** — chance. But the
  metric is INVALID as a role readout: 99/200 predictions echo S2 with agent-before-patient
  order (66 exact copies), and among those textually role-preserving echoes the forced choice
  "tracks" only 23% — cos to "The <noun> was <pp>." probes is dominated by which noun is the
  probe's SUBJECT, not by role (SONAR's surface code again, biting the instrument). Scored
  against prereg as a miss regardless (P4).
- **Judge cell (codex, neutral framing)**: 177 tracks / 15 inverted / 8 unclear →
  **92.2% tracking among determinate — an UPPER bound** (judge saw story+continuation, so
  role-neutral continuations resolve from the story; pre-logged caveat in judge_054.json).
  The clean signal: only **7.8% active role inversions** (e.g. "The fox chased the rabbit." →
  "The rabbit chased the fox."), concentrated in reversible-chase pairs.
- Honest summary: the LCM's next-concept prediction mostly REPEATS/paraphrases the role
  sentence (base-model behavior on 2-sentence prompts), which preserves roles textually; true
  generative role tracking beyond echo is NOT certified by this cell.

## Geometry (lesson-3 profile, geometry_054.json)
Norm mean 0.221 (z) → 64.2 (prenet) → 135.9 (ctx_mid) → 30.1 (ctx_final); anisotropy (mean
pairwise cos) 0.297 → 0.505 → 0.883 → 0.764; prenet has a dominant PC (14.9% EVR vs 5.5% in z).
Any delta/subtraction construction across depths must renormalize.

## Brier vs frozen PREREG_LITE.md predictions
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 z_raw anchor replicates | 0.80 | YES (0.796 ∈ [0.74,0.86]; CI ∋ 0.5) | 0.040 |
| P2 ceiling +0.03 at some depth | 0.30 | NO (max +0.016) | 0.090 |
| P2b ctx_final within ≤ z−0.03 | 0.45 | NO (−0.000 mlp) | 0.203 |
| P3 primary chance at all depths | 0.85 | YES | 0.023 |
| P4 forced-choice ≥0.60, p<0.05 | 0.55 | NO (0.465, p=0.856) | 0.303 |
**Mean Brier = 0.132** (0.114 over the 4 main predictions).

## Limitations / what did NOT run
- Mimir ≠ Meta's checkpoint (weaker training budget than the paper's 1.3T-token models);
  official-LCM claims remain untestable (weights withheld).
- Battery states taken at sequence position 0 with NO preceding context — the battery cells test
  the tower's per-sentence TRANSFORM, not accumulated multi-sentence planning state. (The
  role-swap cell does use 2-sentence contexts.) A position->1 battery variant did not run.
- Forced-choice probe design confounded by surface subject (found post-hoc); judge cell is an
  upper bound by construction. Denoiser-internal states not probed. One diffusion sample per
  context (seed 42). Strict 0.9 battery gate fails for all reps (as for SONAR itself).
- Follow-up worth funding? **Y (narrow)**: (i) battery on states at position k>1 under
  role-relevant multi-sentence contexts (tests planning pressure properly); (ii) role-diagnostic
  probe pairs with novel subjects to fix the forced-choice instrument; (iii) denoiser tower
  depths. As a cheap fold-in, not a new flagship row.

## Ops footprint
Box dir 5.2 GB (venv_lcm ~5 GB + lcm_repo + npz caches; no ckpts of ours); HF cache +3.3 GB
(models--mimir-lcm--Mimir-1.6B, shared). SONAR fairseq2 assets were already cached. Disk 556G
free. tmux c100_054 self-ended; no foreign process/session touched. night8 venv untouched.
