# RESULT — 052 laser-labse (EXT: is SONAR's probe power the MT objective?)

**Tier: T3-exploratory.** Self-harvested in-session 2026-08-03. GPU phys0 (CVD=0) only, guarded.
tmux c100_052; smoke (incl. LASER batching self-check, max err 2.9e-05) passed before full run;
full run wall ~7 min, **0 steps errored**, sentinel DONE.

## Verdict (one line)
**Half the MT hypothesis survives, sharpened: the generative MT-DECODER objective (LASER)
reproduces SONAR's regime — SONAR-band within-construction ceilings + a strong transferable
surface-order code + NO abstract binding — while translation-RANKING (LaBSE, 471M) stays in the
small-contrastive regime despite 4–20x the parameters. It is the decoder, not translation data or
capacity. Bonus inversion: LaBSE hides a lexically-anchored possessor code (s↔of cross 0.986 mlp)
that LASER completely lacks (0.57), and LASER does NOT reproduce SONAR's German case→role binding
(deu primary 0.44 vs SONAR 0.668) — that cell stays SONAR-specific.**

Both models are formally **INSTRUMENT_FAILURE** under the strict 0.9 battery gate — but so is
SONAR itself on stimuli_v2 (amendment A1); the gate-relative comparison below is the science.

## What actually ran
stimuli_v2 battery UNCHANGED (only encoder swapped; 051 machinery): agent_patient + genitive,
linear+mlp, seeds 0,1,2, n_boot 1000, on `laser2` (BiLSTM MT seq2seq, 44.6M enc params, d=1024,
max-pool; laser_encoders 0.0.2, fairseq stubbed — LSTM path is pure torch) and
`sentence-transformers/LaBSE` (BERT dual-encoder translation-ranking, 471.5M, d=768, CLS+dense,
L2-normalized). Secondary pre-registered cell: 061 case-marking battery (vendored run_061_mt.py,
linear, seeds 0,1,2, n_boot 1000) on deu+eng for both models. Artifacts: `out/`
binding_battery_{laser2,LaBSE}.json, BATTERY_RESULTS_052.md, results_061mt_*.{json,md}, run.log.

## Headline numbers — English battery (051/SONAR comparison)
(ap = agent_patient; within = linear/mlp; primary = cross-construction + lexical holdout, best readout)

| model | params | objective | ap within | gen within | ap primary | surf within / cross | gen s↔of cross (lin/mlp) | strict verdict |
|---|---|---|---|---|---|---|---|---|
| SONAR (anchor) | — | MT seq2seq + decoder | 0.758/0.796 | 0.826/0.835 | 0.508 [0.489,0.527] | high (051) | **0.968** (amended pos-ctrl) | INSTR_FAIL → amended NO_BINDING |
| **LASER2** | 45M | MT seq2seq + decoder | **0.799/0.808** | **0.818/0.826** | 0.493 [0.481,0.504] | 0.735 / **0.694** | 0.54 / 0.57 | INSTRUMENT_FAILURE |
| **LaBSE** | 471M | translation ranking | 0.532/0.622 | 0.607/0.702 | 0.500 [0.494,0.505] | 0.529 / 0.538 | 0.61 / **0.986** | INSTRUMENT_FAILURE |
| 051 six | 22–110M | contrastive/retrieval | 0.54–0.60 | 0.61–0.74 | 0.497–0.509 | ~0.52–0.54 | not extracted | INSTRUMENT_FAILURE ×6 |

- **LASER lands IN SONAR's within-ceiling band** (0.80–0.83 vs SONAR 0.76–0.84) with an elevated
  transferable surface code (surf-cross 0.694 vs 051's 0.52–0.54) — the SONAR regime, at 45M
  params and an LSTM. **LaBSE at 471M patterns with the 22–110M contrastive six.**
- **Primary (cross+lex holdout) dead chance for both** (0.49–0.50) — no abstract role binding,
  same as SONAR and all previous encoders. z_bag role = 0.500 everywhere (stimuli sanity ✓).
- **The s↔of inversion (not pre-registered, flagged exploratory):** LaBSE's mlp transfers
  possessor across OPPOSITE-parity genitive constructions at 0.986 when vocabulary is shared
  (the exact cell that certified SONAR's amended positive control, 0.968) while its lexical
  holdout stays 0.538 → a lexically-ANCHORED possessor code (content-addressed, order-invariant
  for seen nouns), not abstract binding. LASER lacks this entirely (0.54–0.57). Reading:
  **generative MT installs a construction-general SURFACE code; translation-ranking installs a
  prop-anchored SEMANTIC code; SONAR has both; abstract binding emerges in neither objective.**

## Secondary cell — 061 case-marking battery (deu + eng control)
| model | lang | verdict | within-role | role cross-lex (PRIMARY) | role cross-all | surf cross | case lexical (pos-ctrl) |
|---|---|---|---|---|---|---|---|
| SONAR (061 anchor) | deu | BINDING_PRESENT | 0.994 | 0.668 [0.608,0.731] | 0.049 | 0.951 | 1.000 |
| **laser2** | deu | **NO_BINDING** | 0.990 | **0.442** [0.386,0.494] | **0.011** | 0.989 | 0.856 [0.824,0.888] |
| laser2 | eng | NO_BINDING | 0.995 | 0.460 [0.402,0.513] | 0.046 | 0.954 | — |
| LaBSE | deu | INSTRUMENT_FAILURE | 0.675 | 0.492 [0.441,0.541] | — | 0.501 | **0.984** [0.975,0.990] |
| LaBSE | eng | INSTRUMENT_FAILURE | 0.616 | 0.479 [0.431,0.528] | — | 0.595 | — |

- **LASER is fully powered on the 061 stimuli (0.99, = SONAR) and reads the German case marker
  (0.856), yet shows NO German case→role binding** — primary 0.442 vs SONAR's 0.668, and extreme
  anti-transfer cross-order (0.011: a pure surface-order reader). SONAR's German binding is NOT
  an MT-objective universal; something SONAR-specific (attention-pooling, capacity, or NLLB data)
  produces it.
- LaBSE deu reads the case marker at 0.984 while its role ceiling is 0.675 — encodes morphology,
  doesn't deploy it for role (null uncertified, power gate failed, per 061 verdict logic).

## Gates / instrument controls
- Strict battery gate (within ≥0.9 AND genitive cross+lex ≥0.9): FAIL both models → formal
  INSTRUMENT_FAILURE, nulls uncertified. Amendment A1 (logged pre-full-run): SONAR itself fails
  this gate on stimuli_v2; SONAR-anchored power = within in 0.76–0.84 band + s↔of ≥0.9. Under A1:
  LASER passes the band leg, fails s↔of; LaBSE the reverse. Neither is fully SONAR-powered.
- z_bag 0.500 ± 0.001 everywhere ✓; laser batching self-check ✓; 061 deu case pos-ctrl CI-lo
  0.824 (laser) / 0.975 (LaBSE) ≥ 0.6 ✓; 0 errors in run.log.

## Brier vs frozen PREREG_LITE.md predictions
- **(a) P=0.65** LASER passes the full 0.9 gate → **FALSE** (ap within 0.808; gen cross+lex
  0.49). Written against the mistaken "SONAR ~1.0" anchor (A1); scored as written. → 0.4225
- **(b) P=0.70** LaBSE fails the gate, patterning with the contrastive six → **TRUE**. → 0.09
- **(c) P=0.75** conditional on LASER powered: SONAR pattern (primary CI-lo <0.6, surface
  elevated) → **antecedent FALSE** (strict gate and A1 both fail) → EXCLUDED. Directionally the
  predicted pattern did hold (primary 0.49, surf-cross 0.694). Sensitivity: Brier would be 0.20
  scoring it TRUE, 0.30 scoring it as-a-miss.
- **(d) P=0.60** ordering LASER > LaBSE > best-051 on ap within (best readout): 0.808 > 0.622 >
  0.604 → **TRUE** (last leg razor-thin, 0.622 vs 0.604 — flagged). → 0.16
- **(e) P=0.55** conditional deu-powered LASER (0.990 ✓, case-ctrl ✓): no German binding
  (CI-lo 0.386 ≤ 0.6) → **TRUE**. → 0.2025
- **Brier (4 scored) = 0.219.** Cost concentrated in (a) — the anchor error, already dissected
  in A1.

## Limitations
- Pooling confound noted in prereg: LASER max-pools, SONAR attention-pools; objective and pooling
  not separable in this pair (LaBSE is CLS+dense — a third pooling family). The clean follow-up
  is a pooling swap on frozen states.
- LaBSE outputs are L2-normalized by the model; laser2 is the 2020 retrain, not 2018 LASER.
  s↔of inversion is post-hoc/exploratory (single cell, though 0.986 with tight CI, both
  directions, all 3 seeds). 061 cell linear-only. One stimuli set. All T3.

## Follow-up worth funding? **Y (narrow).**
(1) Pooling-swap control (attention-pool LASER's states / max-pool SONAR's) to isolate the
decoder-objective vs pooling contribution to the SONAR regime. (2) The LaBSE lexically-anchored
possessor cell (0.986) deserves promotion into the battery as a named cell ("content-addressed
binding") — it is the first non-SONAR encoder to pass the s↔of positive control. (3) German
case-binding is now SONAR-specific: test NLLB-encoder (SONAR's parent) on the 061 battery to
locate where it enters. Row 053 (role-retrieval prompts) proceeds as planned for restoring
small-embedder power.
