# RESULT — row 062 crosslingual-role-transfer

**Status:** DONE + self-harvested in-session. **Tier: T3-exploratory.** Model: SONAR
(`text_sonar_basic_encoder`), linear readout, seeds 0/1/2, n_boot 1000, wall 30 s.
GPU: CVD=0 → phys0 only (guarded, freed to 35 MiB). tmux c100_062 (ended, killed — ours
only). Prereg: adopted unchanged from the t43 manager (killed pre-compute) + one
pre-compute amendment (GPU move, exploratory 4×4 matrix, corpus instantiation).

## Question
Does a role probe trained on English SONAR z transfer zero-shot to translation-parallel
German/Japanese/Turkish (held-out propositions AND vocabulary)? Is 061's weak German
case→role code the same subspace an English probe reads?

## Headline
**Primary transfer cells: INSTRUMENT_FAILURE by the preregistered power gate — and the
failure is precisely diagnosed as a DESIGN flaw, not an encoder fault.** The gate cell
(en-active probe → en-active test-vocab) scored 0.440 [0.367,0.513] because focal-swap
canonicalization + concept-level lexical holdout make ANY surface-order role code
unreadable on unseen nouns by construction: the label "alphabetically-first concept is
agent" requires ranking words the probe never saw. Post-hoc diagnostic: the identical
cell WITHIN vocab (group-CV) is 0.99–1.00 in all four languages (both train-vocab and
test-vocab separately); parallel-sentence retrieval en↔X is P@1 0.88–1.00 (encoder and
row alignment certified); mock-encoder machinery test passed pre-launch. 061's
"within-power ≈0.99" was a within-vocab number; the prereg ported it to a cross-vocab
cell where it provably cannot hold. **Instrument lesson (new, added to the ledger):
power gates must be certified in the same train/test vocab regime as the cell they gate.**

Two genuinely informative yields survive the wreck:
1. **061 replication on fresh vocabulary (native cross-order + lexical cell): Japanese
   REPLICATES AND STRENGTHENS — German FAILS.** jpn 0.696 [0.635,0.751] (061: 0.658,
   inconclusive → now clears the 0.6 bar on a second, disjoint lexicon); deu 0.551
   [0.478,0.624] (061: 0.668, CI-lo 0.608 — gone); tur 0.486, eng 0.521 (both chance,
   as in 061). 061's caveat-1 warning ("German positive is marginal and fragile") was
   right: the deu signal was lexicon-specific. The surviving order-invariant case→role
   binding is JAPANESE (particle-marked), not German (article-marked) — the
   morpheme-separability gradient needs reordering: particle > article ≈ suffix ≈ none.
2. **The one vocab-general role code found (jpn) does NOT transfer crosslingually.**
   jpn-trained order-free probe → eng 0.445 [0.389,0.497], deu 0.450 [0.386,0.512],
   tur 0.451 [0.389,0.510] — slightly BELOW chance, CIs exclude 0.6. Whatever
   order-invariant binding SONAR has in Japanese is language-LOCAL in z, not a shared
   thematic-role subspace. (All other matrix cells ≈ chance too, but those are
   overdetermined by the design flaw; the jpn row is the informative one because jpn
   demonstrably has a vocab-general code to transfer.)

## Numbers (envelope over 3 seeds; full JSON in out/results_062.json)
### en_active probe → target (PRIMARY, all INSTRUMENT_FAILURE via power gate)
| target | congruent (agent-first) | incongruent (patient-first) | pooled |
|---|---|---|---|
| deu | 0.493 [0.420,0.567] | 0.500 [0.424,0.574] | 0.495 [0.469,0.521] |
| jpn | 0.391 [0.322,0.464] | 0.617 [0.540,0.689] | 0.511 [0.479,0.545] |
| tur | 0.587 [0.515,0.659] | 0.424 [0.351,0.502] | 0.509 [0.469,0.550] |
| eng (gate) | **0.440 [0.367,0.513] → FAIL (<0.9)** | 0.533 | 0.485 |

(en_passive mirror: pooled 0.489–0.504 everywhere, chance.)

### Native cross-order + lexical holdout (061 replication, new parallel lexicon)
eng 0.521 [0.446,0.598] · deu **0.551 [0.478,0.624]** · jpn **0.696 [0.633,0.752]** ·
tur 0.486 [0.430,0.543]

### Order-free 4×4 matrix (train row → test col, pooled; diag = native order-free)
All 16 cells in 0.366–0.522; no CI-lo above 0.455. Oddity: jpn→jpn 0.366 [0.303,0.431]
significantly BELOW chance (lexical anti-generalization under both-orders training —
consistent with the vocab-extrapolation pathology, flagged for the redesign).

### Post-hoc diagnostic (within-vocab group-CV; NOT preregistered)
agent-first-family cell: eng 0.998/1.000, deu 0.998/1.000, jpn 0.992/1.000,
tur 0.994/1.000 (train-vocab / test-vocab). Both-orders within-vocab: eng 0.793 <
deu 0.863 < jpn 0.935 < tur 0.967 — within-vocab order-decorrelated role readability
follows the case-marking gradient (Turkish suffix+noun conjunctions are learnable when
the nouns are seen; they just don't extrapolate).

### Controls (instrument lesson 3)
- Language identity is linearly ≈ the mean offset: 4-way lang-ID probe 1.000 raw →
  0.273 (chance) after per-language train-mean centering. Offsets are large:
  ‖μ_L−μ_M‖ / mean‖z‖ = 0.30–0.46.
- Yet mean-centering changed every transfer AUC by exactly 0.0000 — as pre-logged in
  the amendment (probe standardizes by train mean; AUC is rank-invariant to constant
  logit shifts). The centering control is structurally inert for this probe family;
  norm-profile + lang-ID are the real mean-shift diagnostics.
- Retrieval gate: en↔deu 1.00/1.00, en↔jpn 1.00/1.00, en↔tur 0.884/0.992 — PASS.
- Probe-direction cosines (report-only; probes at chance are noise-fitted, do not
  over-read): order-free directions cluster {eng,deu} vs {jpn,tur} with NEGATIVE
  cross-block cosines (e.g. deu·tur −0.86) — direction similarity does not predict
  transfer (cos(jpn,tur)=0.77 yet jpn→tur = 0.451).

## Predictions → Brier (frozen in PREREG_LITE, scored as registered)
| pred | prob | outcome | note |
|---|---|---|---|
| (a) surface code transfers (≥2 targets congruent CI-lo>0.6, incong<0.5) | 0.60 | **FALSE** | unpassable by construction (design flaw); credence was misplaced on an impossible event |
| (b) en→de pooled CI-lo > 0.6 | 0.08 | **FALSE** | 0.495 [0.469,0.521] |
| (c) deu native cross-order+lexical replicates (CI-lo>0.6) | 0.55 | **FALSE** | 0.551, CI-lo 0.478 — 061's German positive did not survive fresh vocab |
| (d) retrieval P@1 ≥ 0.8 all pairs both directions | 0.85 | **TRUE** | min 0.884 |
| (e) centering shifts any en→X pooled by > 0.05 | 0.35 | **FALSE** | all deltas exactly 0.0000 (structural, pre-logged) |

**Brier(a–e) = 0.1628.** Dominant cost = (a): the 0.60 was assigned to a cell that could
not pass under the design — the prereg SHOULD have caught that the cited 0.99 power
figure came from a different vocab regime. (c) was an honest coin-flip on a fragile
positive that landed tails.

## What did NOT run / limitations
- Nothing truncated; all preregistered + amendment cells ran (30 s wall). The post-hoc
  within-vocab diagnostic was added AFTER seeing the gate failure (labeled throughout).
- Because the primary design is flawed, **the crosslingual-transfer question remains
  OPEN, not answered negative** — except the jpn row (informative negative above).
- Linear readout only; one construction pair per language; templated stimuli; one encoder.
- jpn 0.696 is a 2-lexicon replication but still linear-only, template-family-bound, T3.

## Follow-up worth funding? **Y (two narrow threads).**
1. **Transfer-cell redesign that can extrapolate:** score RELATIVE codes instead of
   focal-canonicalized absolutes — e.g. probe on swap-pair differences z(AB)−z(BA)
   (label = which of the two orders is agent-first for the SAME noun pair), which is
   vocab-general by construction; then rerun the crosslingual matrix. This is the fix
   that lets the row's actual question be answered.
2. **Japanese is now the binding thread:** 0.696 across two disjoint lexicons is the
   strongest order-invariant role signal the program has. Breaker next: MLP readout,
   particle-swap null (が/を swapped → signal must invert), more construction types,
   third lexicon. German demoted to cautionary tale.

## Provenance / hygiene
Prereg adopted from killed predecessor + pre-compute amendment (both in PREREG_LITE.md).
Mock machinery test local (uv, all cells chance, gates fired correctly on garbage
encoder). Box: smoke (SMOKE_OK) then full in tmux c100_062, CVD=0→phys0 (guard passed,
freed to 35 MiB); session killed (ours only), no foreign process touched. Box dir 26M
(z_062.npz kept on box, not copied to repo). Disk 570G free. out/: results_062.json,
results_062_table.md, smoke_results.json, mock_results.json, diagnostic_withinvocab.json,
per-lang stimuli + meta, run.log, DONE. Local commit, no push.
