# 093 stimuli-coevolution — RESULT (self-harvested 2026-08-08)

**Tier T3-exploratory · BUILD/adversarial (block L instrument row).**

**Verdict: a codex generate→evaluate→select loop, explicitly tasked to break the agent_patient
binding probe, produced 24 fluent construction variants — 24/24 WELL-FORMED (canon_check PASS),
11 well-formed AND plausible — and NONE broke the probe's integrity. On SONAR, all 11 plausible
families keep z_bag role-blindness at 0.500 (0/11 breaks), move the within-probe by ≤0.05 AUC,
and flip no verdict cell. The only stimuli that break the probe are deliberate structural
confounds (leak-marker / off-by-one / balance-drop), and canon_check catches every one on its
correct property. Breaks are ARTIFACTS the gate defuses, not genuine probe weaknesses —
SCOPED to construction/template attacks (see limitation).**

## Loop design (what ran)
- **Generate (codex LOCAL, neutral framing, stdin prompt, full-output capture, batched):** 3
  rounds × 8 = 24 candidate agent_patient *families*. codex supplies a sentence-frame template
  (slots {S1}=surface-first NP, {S2}=surface-second NP, {VERB}=past verb), which surface slot is
  the agent (parity), and a construction label; round r+1 sees round-r verdicts and mutates.
- **Realize (harness, mechanical):** each family realized over 48 shared baseline propositions ×
  both orders = 96 items, with schema / swap-pairs / balance / A_surface_first / splits
  GUARANTEED by construction. This isolates the adversarial content to the *linguistic template
  + parity*; fillers are the fixed baseline entities.
- **Evaluate each on 3 axes:** well-formedness = `canon_check.certify` (row 091); plausibility =
  codex 1–5 fluency + on-task judge (neutral, batched) + rule checks; probe-stress = local proxy
  (per-token φ leak, surface parity) then **SONAR** (box GPU0, 23 s) for the 11 selected.
- **Select** well-formed ∧ plausible ⇒ **v3 stress set = 11 families** (4 agent=S1, 7 agent=S2 —
  both parities, so the flipped-parity surface control retains power).

## Numbers
| stage | result |
|---|---|
| candidates generated | 24 (3 rounds × 8) |
| **well-formed (canon PASS)** | **24/24** — 0 malformed, max per-token φ = 0.000 throughout |
| plausible (fluency≥3, on-task≥0.75, rules) ∧ WF | **11/24** → v3 stress set |
| SONAR power gate (pooled baseline within-role, ppk.certify d=3) | **PASS** (intrinsic within-role AUC 0.548) |
| baseline mean z_bag role AUC | 0.500 |
| **v3 candidate z_bag role AUC** | range **[0.498, 0.502]**, **breaks (>0.55) = 0/11** |
| v3 candidate within-role Δ vs baseline | **[−0.052, +0.019]** |
| v3 candidate cross-from-baseline role AUC | [0.424, 0.674], **>0.70 count = 0** (no false FALSIFIED) |
| **artifact control** (inject confound → canon_check) | leak→**FAIL e**, off-by-one→**FAIL c**, balance→**FAIL b**, clean→**PASS** |

## What breaks the probe — genuine vs artifact
- **Genuine breaks: NONE found.** No codex-generated, plausible, well-formed stimulus flips a
  SONAR verdict cell. z_bag stays role-blind (lexical leak impossible under the alphabetical-focal
  canonicalization + shared fillers); within-probe barely moves; cross-construction transfer stays
  centered on chance (~0.5, parity-reflected, none >0.7).
- **Artifacts: the only breakers.** A stimulus that inflates/fakes a verdict must inject a
  structural confound (a role-correlated token, a slot-index off-by-one, or a broken swap-pair).
  `canon_check` (row 091) flags each with the correct targeted property → correctly labeled an
  ARTIFACT, not a probe weakness. **The breaker feature is surface/lexical/structural leakage; the
  gate defuses it.**
- **Why template attacks can't win:** the probe label is "the *alphabetically-first* filler holds
  the role," which is orthogonal to any construction/adverb/framing codex can add — so per-token φ
  stays 0 and z_bag stays 0.5 no matter how exotic the fluent frame.

## Predictions & Brier (frozen in PREREG_LITE)
- **P1** codex finds ≥1 candidate that breaks the integrity oracle — P=0.90 → **FALSE** (0/24
  broke it via generation; only manual injection did). Brier 0.81. *Miss: over-anticipated a
  generative break; the canonicalization made template-only structural breaks impossible.*
- **P2** every break FAILS canon_check (artifact) — P=0.60 → **TRUE** (3/3 injected breaks fail on
  the correct property). Brier 0.16.
- **P3 (frozen y/n)** ≥1 plausible+WF candidate flips a SONAR verdict cell — P=0.20 → **FALSE**
  (0/11; z_bag flat, no cell flip). Brier 0.04.
- **P4** dominant breaker feature is surface/lexical leakage, not a genuine cross gap — P=0.70 →
  **TRUE** (the only breakers are the injected surface/lexical/structural confounds). Brier 0.09.
- **P5** high-diversity plausible families leave cross transfer at chance (no false FALSIFIED) —
  P=0.70 → **TRUE** (cross role AUC ~0.5, none >0.7). Brier 0.09.
- **Mean Brier ≈ 0.238** (P1 dominates the loss).

## Limitations (honest)
- **Attack surface scoped to CONSTRUCTIONS/TEMPLATES.** The harness reused the fixed baseline
  entity fillers to guarantee comparability, so codex could vary the *frame + parity* but NOT the
  *filler pools*. The strongest known probe-fragility axes — filler lexical-diversity (row 006) and
  disjoint agent/patient filler pools (an embedding-space z_bag leak that per-token canon_check
  could miss) — were NOT exposed to codex. "No plausible break" therefore holds for template/
  construction evolution; **filler-level adversarial evolution is untested and is the obvious
  follow-up** (row-100 benchmark candidate).
- Probe-stress SONAR run is small-n (48 props/family, single seed, linear readout). Per-family
  within-role AUC (~0.62) is a small-sample proxy and sits below the full battery's within-ceiling;
  the load-bearing quantities are the RELATIVE deltas + the z_bag flatness + the PASS power gate,
  which are stable. No cluster-bootstrap CIs on the per-family cells (point AUCs only).
- codex-as-judge for plausibility is single-model; on-task was strict (it marked baseline objrel
  on-task=0 — a relative-clause quirk), which the fluency+rule floor tolerates.

## v3 stress set (for row 100)
`out/families_v3.json` — 11 well-formed + plausible construction families (spec + realized items),
both parities. Because none breaks the probe, its value to row 100 is as a **hardened
cross-construction generalization battery** (diverse fluent frames that a genuine binding probe
should transfer across), not as a probe-breaker set. Companion: `candidates_table.json` (full
24-family scores), `all_candidates_full.json`, `artifact_control.json`, `sonar_stress.json`.

## Follow-up worth funding? **Y (narrow).** Give codex the filler-pool + lexical-diversity levers
(the 006 axis) as the attack surface and re-run the loop with the full battery + CIs: that is the
one place a plausible+well-formed z_bag/STIMULI_INVALID break might actually exist. Absent that,
the template-level conclusion (probe + canon_check gate robust to adversarial construction
evolution) is a solid T3 instrument-validation result.
