# RESULT — Row 060 human-calibration (BUILD) — T3-LLM-PROXY — CLOSES BLOCK F (051–060)

*Run 2026-08-03, fully LOCAL (no box, no GPU: SONAR z reused from 058's repo cache).
PREREG_LITE.md frozen before the pilot; 5-pair parse smoke per CLI run and deleted
pre-freeze. Panel = codex (gpt-5.6-sol via codex-cli 0.145.0), grok (grok CLI 0.2.118),
agy (Antigravity CLI 1.1.10, Gemini-backed) — each called non-interactively with stdin
closed, neutral framing, full transcripts in `out/raw/` (51 calls, 100% rc=0, zero
retries; wall ≈ 13 min).*

**Tier: T3-LLM-PROXY, hard ceiling.** Nothing in this row is human data. The panel is an
honest *proxy pilot* for the registered human study; the durable artifact is
`STUDY_DESIGN.md` (complete runnable protocol: stimuli, 3 tasks, between-subject
instruction manipulation, N=160 power analysis, exclusions, its own registered
predictions HP1–HP5 to be Brier-scored when humans actually run).

## Question
The 051–059 arc showed pooled embeddings don't bind roles and their similarity is
topic/surface-driven. Is that a deviation from HUMAN similarity judgment, or do casual
human judgments also let topic dominate roles? This row (a) designs the definitive human
study and (b) pilots it on an LLM panel: do LLM judges rate role-swap pairs ≈ paraphrase
pairs (topic-dominance) or sharply lower (role-sensitivity)? Does a role-critical
instruction move them? Is SONAR z-cosine closer to "casual" or "role-critical" judgments?

## What ran
- `src/build_pairs.py` (seed 60): 100 pairs, 5 cells × 20, every proposition in exactly
  one cell, all sentences from the frozen stimuli_v2 agent-patient battery:
  SURF (active/cleft, same meaning, minimal surface), PARA (active/passive, same meaning,
  order flipped), RS-SAME (active AB/BA role swap, identical words), RS-CROSS
  (active-AB/passive-BA: meaning flips, surface entity order PRESERVED), TOPIC
  (different entities+verb). Plus 30 forced-choice triplets (anchor active-AB; options
  passive-AB paraphrase vs active-BA role swap, sides randomized).
- Panel: 3 models × 3 rating conditions (casual_v1, casual_v2 wording-variant
  pseudo-panelist, role_critical) × 100 pairs, batches of 20, opaque shuffled ids +
  3 × 30 triplets ("which option means the same thing as X?").
- Calibration: SONAR z-cos per pair from 058's `z_text.npz` (G4: 0/230 cache misses),
  content-word Jaccard baseline, Spearman panel-vs-z, bootstrap CI on the difference.
- `src/analyze_060.py` → `out/results_060.json`.

## Gates — all PASS (the panel is a valid instrument)
- **G1 parse**: 100% valid ratings in all 9 panelist×condition cells and all 3 triplet
  sets (1230/1230 judgments; zero retries needed).
- **G2 panel consistency**: pairwise Spearman (casual_v1) codex–grok 0.972, codex–agy
  0.962, grok–agy 0.942 → mean **0.959** (floor was 0.30). Wording-variant
  pseudo-panelists replicate their parent model at ρ = 0.955–0.971.
- **G3 anchor sanity**: every model SURF ≥ 6.85, TOPIC ≤ 1.05.
- **G4 z-cache**: 0/230 sentences missing.

## Panel results (cell means, 1–7 scale; pooled over 3 models)

| cell | casual (v1+v2) | role-critical | SONAR z-cos |
|---|---|---|---|
| PARA (same meaning, order flipped) | **6.98** | 7.00 | 0.761 (lowest same-topic!) |
| SURF (same meaning, minimal change) | 6.93 | 7.00 | 0.910 |
| RS-CROSS (meaning flipped, order preserved) | 3.03 | 2.60 | **0.915** (highest!) |
| RS-SAME (meaning flipped, words identical) | 2.94 | 2.55 | 0.808 |
| TOPIC (unrelated) | 1.05 | 1.05 | 0.373 |

- **P1 TRUE (far beyond threshold): the panel is hard role-sensitive even under casual
  framing** — casual PARA − RS-SAME gap = **4.04** points (predicted ≥ 1.0). No model
  shows topic-dominance (per-model casual_v1 gaps: codex 3.7, grok 4.5, agy 3.65).
- **P2 FALSE (informative miss)**: the role-critical instruction widens the gap by only
  +0.41 (predicted ≥ +1.0) — a **ceiling effect**: the panel already reads for meaning
  under casual instructions, so there is almost no casual regime to move away from
  (item-level Spearman between the casual and role-critical panel matrices = **0.96**).
  Grok is the most instruction-sensitive (RS cells 2.60 → 1.70).
- **P3 TRUE**: topic similarity is still honored — RS-SAME (2.94) sits 1.89 above
  TOPIC (1.05); role-swaps are "same topic, different event", not "unrelated".
- **P4 TRUE**: z-cos inverts the meaning ordering (RS-SAME 0.808 > PARA 0.761;
  RS-CROSS 0.915 ≈ SURF 0.910).
- **P5 FALSE (informative miss)**: z-cos is NOT closer to the casual panel —
  Spearman(z, casual) = 0.379 vs Spearman(z, role-critical) = 0.425 (diff −0.045,
  bootstrap CI [−0.097, +0.002]). The premise failed: the panel has no distinct
  "casual" similarity regime for z to match.
- **P6 TRUE, maximally**: explicit-meaning triplets — panel picks the paraphrase
  **90/90**; SONAR z-cos picks the role-swap **30/30** (mean margin +0.043, min
  +0.006). A perfect judge/embedding inversion.

**Brier (P1–P6) = 0.233** (0.0625 / 0.64 / 0.09 / 0.16 / 0.4225 / 0.0225). Both
misses are the same lesson: I over-estimated how "casual" an LLM panel can be.

## The z-side facts (deterministic, computed at build time)
| cell | meaning | surface entity order | z-cos mean (sd) | content-word Jaccard |
|---|---|---|---|---|
| SURF | same | preserved | 0.910 (0.017) | 1.00 |
| RS-CROSS | **flipped** | preserved | **0.915** (0.016) | 1.00 |
| RS-SAME | **flipped** | flipped | 0.808 (0.035) | 1.00 |
| PARA | same | flipped | 0.761 (0.040) | 1.00 |
| TOPIC | different | — | 0.373 (0.064) | 0.01 |

SONAR z-cos groups the four same-topic cells purely by SURFACE ENTITY ORDER, not
meaning: {order-preserved: RS-CROSS 0.915 ≈ SURF 0.910} ≫ {order-flipped: RS-SAME
0.808 > PARA 0.761} — the role-swap-with-preserved-order is judged (by cosine) as
similar as a near-verbatim restatement, and the true paraphrase is the LEAST similar
same-topic pair. In the triplets, z-cos prefers the role-swap to the paraphrase
**30/30** (mean margin +0.043, min +0.006).

## Calibration verdict (LLM-PROXY tier)
**SONAR z-cosine matches NEITHER panel regime — it is anti-meaning within topic.**
Restricted to the 80 same-topic pairs (where content-word Jaccard is 1.00 by
construction and cannot help), Spearman(z, panel-casual) = **−0.218**: the more
similar SONAR thinks two same-topic sentences are, the *less* similar the panel judges
them, because z's within-topic variance tracks surface entity order and the panel's
tracks who-did-what. Overall a bag-of-content-words Jaccard predicts the panel BETTER
than SONAR does (ρ 0.707 vs 0.379) — the embedding's extra within-topic structure is
not noise relative to judged meaning, it is *anti-signal*. So at proxy tier, the arc's
"failure" framing survives its first judgment-side test: pooled-embedding similarity
deviates from every meaning-attentive judge we could instantiate, and the deviation is
sign-inverted, not just attenuated.

**The honest caveat that keeps the human study necessary**: the pilot FAILED to
instantiate a casual topic-dominant judge. Frontier LLMs read for meaning under any
instruction (casual ≈ role-critical, ρ = 0.96) — so this pilot cannot rule out that
*casual humans* (skimming, gist-level) show the topic-dominance regime that would
partially vindicate embedding behavior. What it does establish: (i) an LLM panel
cannot serve as a proxy for casual human raters (methods finding — LLM "crowd
substitution" papers take note); (ii) if humans pattern like ANY attentive judge, the
gap is real and large (≈4 points on a 7-point scale); (iii) the instrument (items,
batching, tasks, exclusion logic) transfers to humans unchanged.

## Limitations / what did NOT run
- **No humans.** Every panel number is an LLM-proxy; LLM similarity judgments may be
  systematically more meaning-weighted than casual human raters (they read carefully by
  construction) — the pilot most plausibly OVERSTATES human role-sensitivity, which is
  exactly why HP1 stays at .70 in the design and the human study remains necessary.
- Panel models are 3 correlated frontier assistants (+1 wording variant each), not
  independent raters; consistency numbers are not human inter-rater reliability.
- Single stimulus register (stimuli_v2 templates, symmetric-plausibility court/social
  predicates); no naturalistic sentences; English only.
- Calibration uses SONAR only (the arc's reference encoder); other embedders' cosines
  can be added offline to the same frozen pairs.
- The role-critical instruction is one operationalization of "meaning matters"; the
  human study's T2 triplet task bounds the comprehension ceiling separately.

## BLOCK F SYNTHESIS (051–060): the beyond-SONAR story
Ten cells asked whether ANY pooled representation binds thematic roles, and every one
answered no: six small contrastive embedders barely encode role even within-construction
(051); 44× scale inside the retrieval-contrastive family moves nothing, 110M → 4.8B
(059); instruction conditioning (053), cross-encoder cross-attention (056 — actively
below chance on parity-flipped pairs), latent chain-of-thought feedback (055), discrete-
diffusion denoising at every noise level (057), and an LCM context tower (054) all fail
to install it. The one real regime shift in the block is objective, not capacity: a
generative MT decoder (LASER 45M, 052; SONAR; speech-SONAR 058, which is the same code
with a small removable acoustic accent) buys a high within-construction ceiling and a
*transferable surface code* — while translation-ranking (LaBSE) instead buys a lexically
anchored possessor code — but abstract who-did-what binding emerges in neither regime.
060 closes the block by testing the framing itself: an LLM judge panel (honest proxy for
the registered human study) sharply separates role-swaps from paraphrases (gap 4.0 on a
7-point scale) under any instruction, whereas SONAR's similarity is sign-INVERTED on the
same items (prefers the role-swap over the paraphrase 30/30, within-topic ρ = −0.22
against the panel) — so pooled-embedding similarity is not "matching human coarse
similarity", it deviates from every meaning-attentive judge tested, pending only the
casual-human anchor that STUDY_DESIGN.md is built to nail down. The block's fundable
question is now sharp: what training signal WOULD install binding — and the judgment
anchor for evaluating any candidate is designed, costed (~$1k), and frozen.

## Follow-up worth funding? **Y — the human study itself.**
STUDY_DESIGN.md is runnable as written (~$1,000, N=200 recruit / 160 analyzed,
Prolific + jsPsych). It is the single missing anchor for the whole 051–060 arc: it
decides whether "embeddings don't bind roles" is a deviation from human similarity or a
faithful model of its casual regime. Everything else in the arc is now bottlenecked on
that anchor, not on more encoder cells.
