# 001 roleswap-contrastive — RESULT (harvested 2026-08-01, tier T3-exploratory)

**Verdict: NEGATIVE — the contrastive role-swap objective was satisfied WITHOUT inducing
a transferable role code.**

## Numbers
- Training completed full budget (5e8 tgt tokens, 43,898 steps, val_f1 0.6444 — no
  reconstruction cost vs baseline's 0.6554 beyond ~0.011).
- **InfoNCE succeeded at its own task**: 1.33 → 0.32 final (anchor ranks positive over the
  role-swapped hard negative with ~0.73 softmax prob on training vocab).
- **Battery primary cell (cross-construction + lexical holdout), agent_patient**:
  organism **0.501 [0.486, 0.516]** vs baseline A_D_s0 **0.501 [0.487, 0.516]** —
  literally indistinguishable. MLP readout same (0.492 vs 0.497).
- Within-construction ceiling: organism 0.702/0.777 (lin/mlp) vs baseline 0.664/0.767 —
  at most a marginal within-family shift, nothing that transfers.
- Battery mechanical verdict INSTRUMENT_FAILURE as pre-flagged (organism trains no
  genitive → positive-control gate fails by construction); primary AUC read per prereg.

## Interpretation (scoped)
The organism separated role-swapped pairs *on the training distribution* (InfoNCE ↓)
while showing ZERO gain in construction-invariant role decoding on held-out vocabulary —
i.e., it satisfied the contrast through lexically-specific / idiosyncratic directions
rather than a general role code. λ=0.3 role-swap hard negatives are therefore NOT
sufficient to induce binding in a rung-A DAE organism at this budget. Shortcut
satisfaction of a binding-flavored objective is itself a noteworthy negative for the
"just add a contrastive term" intuition (bears on MANIFEST 096 toy-theory predictions).

## Brier (prereg-lite)
P1 primary CI-lo>0.6 (0.45) → F; P2 val_f1≥0.605 (0.70) → T; P3 organism beats baseline
by ≥0.05 (0.60) → F. Mean Brier **0.218**.

## Caveats / follow-ups
- Confound (pre-logged): differs from baseline in both InfoNCE AND 20% transitive mix-in
  — moot for the headline since the effect is exactly zero, but a λ=0 same-mix control is
  still the clean design if promoted.
- Single λ, single rung, single seed; genitive gate inapplicable to organisms (campaign
  should consider an organism-appropriate positive-control task if promoted).
- Follow-up worth funding? **Yes, conditionally** — 002's dose-response (running) is the
  direct next probe; if λ=1 aux-QA also fails to transfer, the "shortcut satisfaction"
  story strengthens and 096 (toy theory) should formalize it.
