# RESULT — row 061 case-marking-langs (opens MULTILINGUAL block)

**Status:** DONE + self-harvested in-session. **Tier: T3-exploratory.** Model: SONAR
(`text_sonar_basic_encoder`), readout linear, 3 seeds, n_boot 1000, wall 51 s. GPU3 =
CVD=2 → phys3 only (guarded; freed to 15 MiB after). tmux c100_061 (killed, ours only).

## Question
SONAR's English binding battery finds NO transferable thematic-role code (role rides
surface order). SONAR is multilingual. In case-marking languages role is marked by
morphology, not order, so role ⊥ surface-order naturally — the single most promising place
to find binding. **Does SONAR bind role in Turkish / Japanese / German, or is the null
universal across morphology?**

## Headline
★ **NOT a universal null — a morpheme-separability GRADIENT.** SONAR encodes the case
marker at ceiling in every language (case-morphology probe 0.99–1.00), still defaults to a
**surface-order role code everywhere** (the memorization-friendly cross-order cell
`role_cross_all` ANTI-transfers universally, 0.02–0.21 ≪ 0.5), but in the pre-registered
primary cell (cross-word-order + lexical holdout) a **weak order-invariant case→role
binding peaks above chance in German** (0.668, CI-lo 0.608 > 0.6 → BINDING_PRESENT) and
**nearly so in Japanese** (0.658, CI-lo 0.594, INCONCLUSIVE), while **Turkish** (0.413) and
the **English control** (0.520) show none. The order German > Japanese > Turkish > English
tracks how SEPARABLE the case morpheme is as a token: free-standing article `der/den` and
particle `が/を` are read order-invariantly; the bound Turkish suffix `-I` is not, despite
being fully readable as morphology. **The "cheapest big surprise" partially materialises —
marginal, linear-only, T3; a breaker candidate, not a claim.**

## Per-language results
| lang | verdict | within-role (power) | **role cross-lex (PRIMARY)** | role cross-all | surface cross | case lexical (pos-ctrl) |
|---|---|---|---|---|---|---|
| tur `tur_Latn` | NO_BINDING | 0.992 | **0.413** [0.357,0.468] | 0.208 [0.181,0.238] | 0.792 | 0.995 [0.990,0.998] |
| jpn `jpn_Jpan` | INCONCLUSIVE | 0.992 | **0.658** [0.594,0.723] | 0.176 [0.153,0.199] | 0.824 | 1.000 |
| deu `deu_Latn` | **BINDING_PRESENT** | 0.994 | **0.668** [0.608,0.731] | 0.049 [0.037,0.062] | 0.951 | 1.000 |
| eng `eng_Latn` (ctrl) | NO_BINDING | 0.987 | **0.520** [0.454,0.591] | 0.024 [0.016,0.032] | 0.976 | — (no case) |

(within-role & surface_within ~0.97–1.00 all langs; per-seed primary AUCs identical across
seeds — the linear cross-lexical point estimate is deterministic given the frozen split.)

## What each cell means
- **within-role ceiling ≈ 0.99 everywhere** → probe fully powered; role info present when
  position AND case both cue it. (gate PASS all langs; pred e.)
- **case-morphology positive control 0.99–1.00** → SONAR linearly encodes the case marker
  (accusative `-I` / `を` / `den`) with position & lexeme held non-predictive. The nulls are
  therefore INFORMATIVE: z HAS the marker but (mostly) does not BIND it to the entity
  order-invariantly. (pred c.)
- **`role_cross_all` ANTI-transfer universally (0.02–0.21)** → the memorization-friendly
  cross-order probe latches onto surface position, which FLIPS between the two orders → far
  below chance. Anti-strength tracks surface encoding: eng 0.024 (surf 0.98) < deu 0.049
  (0.95) < jpn 0.176 (0.82) < tur 0.208 (0.79). This is "role rides surface order," now
  shown in 4 languages and 3 morphologies.
- **`role_cross_lex` (PRIMARY, disjoint vocab + cross-order)** → the only order-invariant,
  vocab-general signal is the case marker. German clears 0.6, Japanese ties, Turkish/English
  do not. A pure surface reader cannot exceed chance here (all pairs are opposite-parity, so
  pooled == flipped-parity), so deu/jpn > 0.6-ish is genuine (weak) case→role binding.

## Interpretation (the gradient)
Order-invariant case→role readability tracks **morpheme separability**: SONAR's subword/MT
encoding represents a distinct FUNCTION WORD (German article der/den; Japanese particle
が/を) in an order-invariant way well enough for a linear role code, but a BOUND SUFFIX
(Turkish accusative -I fused to the noun) is read only as surface morphology tied to
position, not bound to the entity across orders. English (no case) is the floor. So binding,
where it appears, is a property of *how* the language exposes the marker, not of "case" per se.

## Caveats (load-bearing)
1. **German positive is MARGINAL and fragile.** CI-lo 0.608 barely clears 0.6; linear-only;
   and the companion `role_cross_all` cell is strongly anti (0.049) — the two cross-order
   cells DISAGREE IN SIGN. The binding signal is cell-dependent and weak.
2. **German vs Japanese is a threshold artifact.** 0.668 (CI-lo 0.608) vs 0.658 (CI-lo
   0.594): both are the same weak-positive tier; the BINDING/INCONCLUSIVE split is 0.014 of
   CI-lo. Read them together, not as a qualitative German-only effect.
3. **cross_all vs cross_lexical tension.** The harder, disjoint-vocab cell gives the HIGHER
   role AUC — counterintuitive; likely a regularization/generalization effect (absent lexical
   memorization the L2 probe leans on the weak case signal instead of the dominant, anti-
   transferring surface feature). Needs a breaker to rule out an artifact.
4. Single readout (linear; MLP not run — would strengthen or kill the deu/jpn positive).
   One construction pair per language (2 orders); more orders (OVS, verb-initial, ditransitive
   + dative) would test robustness. Greedy templated stimuli, one encoder.

## Predictions → Brier (frozen in PREREG_LITE)
| pred | prob | outcome | note |
|---|---|---|---|
| (a) ≥1 case lang primary CI-lo>0.6 | 0.45 | **TRUE** | German CI-lo 0.608 (per prereg rule) |
| (b) surface-cross lower in case langs than Eng | 0.55 | **TRUE** | case-mean 0.856 < eng 0.976; each < eng |
| (c) case readable even where role isn't | 0.85 | **TRUE** | case 0.99–1.0 while Turkish role 0.41/anti |
| (d) no lang binds — negative universal | 0.48 | **FALSE** | German binds (marginal) per prereg |
| (e) within ceiling ≥0.75 all 4 langs | 0.90 | **TRUE** | all ≈0.99 |

**Brier(a–d) = 0.1895; Brier(a–e) = 0.1536.** Dominant cost = the near-complementary (a)/(d)
pair: I hedged ~50/50 leaning toward the universal null, and the answer landed (marginally)
on the binding side. Honest note: German's positive is fragile enough (0.608, linear-only,
cross_all disagrees) that a strict/breaker reading could flip (a)→FALSE, (d)→TRUE and LOWER
the Brier — the prereg rule scores it as binding, and I report that.

## Gates
- Surface ⊥ role by construction: generator asserts P(agent-first)=0.500, corr(role,surface)
  =0.000, families 750/750 — PASS all 4 langs (verified pre-encode).
- Power gate within-role ≥0.75: PASS all (≈0.99).
- Case positive control ≥0.6 (case langs): PASS all (0.99–1.00) → nulls informative.
- Mock-encoder machinery test (local): all cells ≈0.5 → INSTRUMENT_FAILURE guard fires ✓.
- On-box SONAR smoke (5 stim/lang): all 4 source_langs encode dim=1024, probes run ✓.
- z-norm 0.19–0.22 all langs (≈ campaign SONAR norm 0.20–0.22; encode sane).

## Example stimuli (both orders = same meaning; role by case, not order)
- **tur** SOV `Garson doktoru itti.` / OSV `Doktoru garson itti.` — "the waiter pushed the doctor"
- **jpn** SOV `画家が少女を見た。` / OSV `少女を画家が見た。` — "the painter saw the girl"
- **deu** SVO `Der Bruder sah den Koch.` / OVS `Den Koch sah der Bruder.` — "the brother saw the cook"
- **eng** active `The doctor chased the lawyer.` / passive `The lawyer was chased by the doctor.`

## Follow-up worth funding? **Y (narrow, breaker-first).**
1. **Breaker on German** before any promotion: MLP readout; more orders (German OVS variants,
   scrambled subordinate clauses); does the signal survive a null where case markers are
   swapped to break der↔den? Is 0.668 an artifact of the article being a separable adjacent
   token (bag-of-words "den+Noun")? → feeds 062 (crosslingual-role-transfer), 073.
2. **Morpheme-separability hypothesis** as a first-class claim: predict binding strength from
   how free-standing the case exponent is (article > clitic/particle > agglutinative suffix >
   fusional). Test agglutinative (Turkish, Finnish) vs particle (Japanese, Korean) vs article
   (German) vs none (English) — is the gradient reproducible? This is the interesting thread.
3. cross_all-vs-cross_lexical sign flip: diagnose whether disjoint-vocab regularization is
   what surfaces the weak case signal (feeds the probe-power lessons).

## Provenance / hygiene
CVD=2→phys3 only (hard-guarded, freed to 15 MiB); venv night8; unset CONDA_PREFIX;
HF_HOME=/workspace/hfcache; BLAS≤8; box dir 2.8M (disk fine, 572G free). Siblings
c100_040/c100_041 (phys0/phys2) + all foreign tmux untouched; never killed a foreign job.
Stimuli generator + runner (binding_battery vendored) in src/; out/{results_061.json,
results_061_table.md, run.log, <lang>/stimuli_agent_patient.json, meta.json}. NOT left
RUNNING. Local commit, no push.
