# RESULT — Row 059 binding-vs-scale (EXT) — T3-exploratory

*Run 2026-08-03, tmux c100_059, launched 15:28:02 UTC, DONE 15:38 UTC. All FOUR scales ran —
the xxl 4.8B fp16 stretch attempt SUCCEEDED on the 16GB A4000 (run.log ends "[059] XXL:
SUCCEEDED / core_fails=0"). Battery walls: base 1.6 / large 1.9 / xl 2.4 / xxl 4.3 min
(total battery ~10.2 min; the rest was HF downloads). Smoke passed pre-launch, smoke
artifacts deleted (051 lesson). Harvested by dedicated agent from out/ artifacts + run.log.*

## What ran
The stimuli_v2 binding battery **UNCHANGED** (051/052 machinery verbatim: agent_patient 2000 +
genitive 1200, seeds 0,1,2, linear+mlp readouts, n_boot 1000, cluster-bootstrap CIs) over the
full GTR-T5 scale ladder — one fixed architecture (T5-encoder retrieval dual-encoder), one
fixed objective (contrastive retrieval), output dim 768 at every scale:

| model | params | dtype | wall |
|---|---|---|---|
| gtr-t5-base | 110,218,368 | fp32 | 97 s |
| gtr-t5-large | 335,726,080 | fp32 | 115 s |
| gtr-t5-xl | 1,241,696,256 | fp16 | 144 s |
| gtr-t5-xxl | 4,865,577,984 | fp16 | 259 s |

xxl was registered as a non-fatal stretch; it fit fp16 on one A4000, so the ladder is 44x in
parameters end to end (110M → 4.8B). Nothing was truncated; 0 errors.

## Gates
- **Probe power** (ap within ≥ 0.9 AND genitive cross+lexical ≥ 0.9): **FAILED at all four
  scales, both readouts** → formal **INSTRUMENT_FAILURE ×4** — exactly the pre-registered
  expectation (arc prior); per 052 A1 measured-anchor logic the scale TREND vs the in-run
  gtr-t5-base anchor is the readout, and every trend below is cleanly measured.
- **z_bag** role AUC = **0.500 exactly** in every cell of every model (label not lexically
  recoverable; stimuli sane).
- Smoke gate: passed pre-launch (base, 100-item subset, seed 0, nb200, gates populated).

## Headline: 44x scale moves NOTHING. Ninth cell of the arc, ninth NO —
## the binding/probe-power picture is scale-invariant within the retrieval-contrastive family

**Scale curves** (mlp readout; aggregate across seeds 0,1,2; CIs are cluster-bootstrap 95%):

| cell | base 110M | large 335M | xl 1.24B | xxl 4.8B |
|---|---|---|---|---|
| ap within (ceiling, KEY curve) | 0.601 | 0.602 | 0.604 | 0.591 |
| genitive within (ceiling) | 0.716 | 0.724 | 0.714 | 0.706 |
| **primary binding** (ap cross+lex pooled) | 0.508 [0.495,0.522] | 0.508 [0.491,0.526] | 0.505 [0.484,0.527] | 0.506 [0.487,0.522] |
| flipped-parity (ap cross+lex) | 0.505 [0.488,0.524] | 0.501 [0.480,0.521] | 0.502 [0.479,0.523] | 0.499 [0.475,0.520] |
| surface-cross (ap, linear) | 0.544 [0.534,0.555] | 0.542 [0.533,0.553] | 0.542 [0.534,0.551] | 0.538 [0.528,0.549] |
| **s↔of possessor** (s_gen→of_gen cross+lex, mlp, 3-seed mean) | 0.489 | 0.472 | 0.455 | 0.542 |

- **Within-construction ceiling: dead flat.** Δ(xl − base) = **+0.003** (0.604 vs 0.601);
  xxl actually dips to 0.591. All four sit inside the 051 small-contrastive / LaBSE band
  (0.54–0.74), nowhere near the 0.9 gate. Capacity within a retrieval-contrastive family
  does not buy probe power — the 052 conclusion ("it's the objective, not capacity") now
  holds across a 44x within-family ladder, not just across families.
- **Primary binding: chance at every scale.** Max CI-lo across all four scales and both
  readouts = 0.496 (base linear). No abstract who-did-what binding emerges anywhere between
  110M and 4.8B. The **9th consecutive NO** of the 051–059 arc.
- **Surface-cross: flat at ~0.54**, slightly declining (0.544 → 0.538). Never approaches
  the SONAR/LASER high transferable-surface regime (<0.60 throughout) — scale does not
  install the generative surface code either.
- **s↔of possessor code: absent at every scale.** The LaBSE 0.986-mlp "content-addressed
  possessor" cell reads 0.489 / 0.472 / 0.455 / 0.542 here (reverse direction of→s: 0.484 /
  0.508 / 0.457 / 0.441 — the xxl 0.542 point-estimate is one-directional, its per-seed
  CI-los graze 0.48–0.50, and the trend is non-monotone). GTR at 4.8B has none of what
  LaBSE has at 471M: the possessor code is a ranking/matching-objective property, full stop.
- **The one cell that moves (exploratory, not prereg'd):** genitive cross-construction
  ALL-VOCAB (lexical shortcut available, mlp): 0.722 → 0.786 → 0.805 → 0.799. Scale does
  improve lexically-mediated transfer — but its lexical-holdout twin stays at chance
  (0.490/0.496/0.473/0.500), so what grows is shared-vocabulary code, not structure. This
  is a nice internal control: the battery CAN see a scale trend when one exists; the
  binding cells' flatness is not an insensitive instrument.

**Honest verdict:** does scale move ANYTHING? On every pre-registered cell, **no** — the
curves are flat to within ±0.02 over 44x parameters, and the only movement (lexical-shortcut
transfer) is precisely the kind of code the arc already knew these models carry. All four
scales are formal INSTRUMENT_FAILURE on the absolute gate, as predicted; as a trend
experiment it is a clean, fully-measured flat result.

## Brier (P1–P4 exactly as frozen; scored on the core base/large/xl ladder)
| pred | P | outcome | check | Brier |
|---|---|---|---|---|
| P1 ceiling flat | 0.75 | **TRUE** | xl within-mlp 0.604 < 0.75 AND Δ(xl−base) = 0.003 < 0.10; all three INSTRUMENT_FAILURE | 0.0625 |
| P2 binding chance | 0.88 | **TRUE** | primary CI-lo ≤ 0.55 at all 3 scales (max 0.496) | 0.0144 |
| P3 no surface regime | 0.70 | **TRUE** | xl surf-cross 0.542 < 0.60 | 0.0900 |
| P4 no possessor code | 0.58 | **TRUE** | xl s↔of mlp 0.455 < 0.70; directional alternative NOT triggered (no monotone growth) | 0.1764 |

**Mean Brier = 0.0858** (4/4 correct). **xxl bonus data (not scored, as registered):**
strengthens every conclusion — P1: xxl within 0.591 < 0.75, Δ(xxl−base) = −0.010 (the curve
bends DOWN, not up); P2: xxl primary 0.506, CI-lo 0.487; P3: xxl surf-cross 0.538; P4: xxl
s→of 0.542 with of→s 0.441 — noise around chance, not emergence, and it breaks any monotone
story. Extending the ladder 4x beyond the registered endpoint changes nothing.

## Place in the 051–059 arc (ninth cell)
051: six small (22–110M) contrastive embedders — all INSTRUMENT_FAILURE, binding chance.
052: LaBSE 471M (ranking) stays in the small band → capacity alone doesn't do it; LASER2
45M (MT seq2seq) lands in the SONAR band → the generative-MT decoder buys the regime.
053–057: instruction-tuned embedders, LCM, COCONUT, cross-encoder, diffusion latents — NO
at every stop. 058: the null is modality-general (speech-z = text-z, same chance profile).
**059 closes the remaining axis — scale.** Held architecture+objective fixed and swept 110M
→ 4.8B: nothing moves. The arc's claim is now fully triangulated: the absence of abstract
role binding in single-vector sentence embeddings is **architecture-general, modality-
general, and scale-invariant; the SONAR-regime exception is bought by the generative
decoder objective, not by capacity** — with the LaBSE possessor cell showing ranking
objectives buy a different (content-addressed) code that scale also does not buy.

## Limitations
- Ceiling honesty as pre-registered: off-the-shelf default encoding (no role-retrieval
  prompts), one stimuli set, linear + 1-hidden-MLP probes only; negative/flat is
  T3-exploratory; no absolute-gate null certified without a powered positive control.
- **Fixed 768-d bottleneck:** the GTR ladder scales the ENCODER (110M→4.8B) but projects
  every scale to d=768 — so "scale" here means encoder compute/depth, not representation
  width. A wider-embedding ladder could in principle behave differently (no such
  off-the-shelf single-family ladder exists).
- Mixed dtype (fp32 base/large, fp16 xl/xxl) — a confound in principle, immaterial in
  practice given curves flat to ±0.02 and fp16 xl reproducing base within 0.003.
- s↔of cell aggregated as 3-seed mean of per-seed pair AUCs (the battery does not pool the
  pair across seeds); per-seed CIs reported in text.
- Nothing failed to run; nothing truncated.

## Tier & follow-up
**T3-exploratory** (campaign default; EXT row). **Follow-up as its own row: N.** The
capacity axis is now closed with four measured points and every cell flat; the marginal
next dollar goes to the arc-level synthesis (and the already-queued 061/062 case-marking /
cross-lingual rows), not a fifth scale point. The one loose thread worth a fold-in, not a
row: the growing genitive all-vocab curve (0.72→0.80) as a case study of what scale DOES
buy retrieval embedders (lexical shortcut sharpening).

## Artifacts
- `out/binding_battery_gtr-t5-{base,large,xl,xxl}.json` — full per-seed batteries
- `out/BATTERY_RESULTS_059.md` — per-model battery tables (box-generated)
- `out/run.log` — full run log (smoke ref, downloads, 4 batteries, sentinel)
- Box: `~/campaign100/059-binding-vs-scale/` (out/DONE sentinel; emb caches ~32 MB×4 left
  on box, not copied); HF cache +~17 GB (gtr base/large/xl/xxl)
