# RESULT — 073 order-code-archaeology: WHEN does the surface-order / anti-transfer code form?

**Tier: T3-exploratory. DONE + SELF-HARVESTED** (tmux `c100_073`, CPU-only CVD="",
launched 08:39:57Z, DONE 08:55:06Z, ~15 min). 18 ckpt-batteries over ladder rung-A
milestone sequences: **A_D** (plain BART-DAE — the arm where 004/005 characterized the
anti-transfer) s0/s1 (4 milestones each) + **A_P90** (paraphrase arm, denser 5-milestone
tail) s0/s1. Battery = stimuli_v2 agent_patient (2000 items, 5 families, P(A surface-first)
=0.5), linear+mlp, seeds 0,1,2, n_boot 1000, same stimuli 001 used. **All gates PASS**
(G1 positive control, G3 z_bag=0.500 every ckpt, G4 norms finite). **Mean Brier 0.276**
(P1 F, P2 T, P3 F, P4 F, P5 T) — the 048-derived "early/front-loaded" prior was **wrong**;
the order code forms mid-training, later than dictionary atoms.

## Verdict (headline — answers the 4 sub-questions)

**The surface-order code and its ANTI-transfer signature form GRADUALLY in MID-training,
are LOCKED to each other (same time, not surface-first-then-anti-later), and track
reconstruction capability. They are NOT an early cheap heuristic and NOT a late abrupt
phase transition.**

1. **When — early or late?** MID. Onset ~step 4000 (**11% of budget**); essentially
   complete by ~step 26000 (**74%**). At step 2000 (5-6% budget) the surface code is still
   at chance (surfX ≈ 0.52). This is **later** than 048's dictionary atoms (~77% crystallized
   by 11%): basic reconstruction (bag-F1 ≈ 0.4) is learned first, order/position-packing
   into z comes after. The "cheap first-learned heuristic" hypothesis is **rejected**.
2. **Abrupt or gradual?** GRADUAL — smooth monotone rise across the mid window, no
   adjacent-milestone jump; A_P90's 5-point grid shows steady intermediate values
   (surfX 0.53→0.56→0.69→0.69) already at final by 74%. (A_D's 11%→100% grid is too coarse
   to resolve the tail alone — the gradual claim rests on A_P90 + cross-arm agreement.)
3. **Anti-transfer at the same time as the surface code, or later?** SAME TIME / LOCKED.
   surfX rising and role `cross_all_flipped` dropping are equal-and-opposite at every
   milestone; both onset ~step 4000 and saturate ~step 26000. The anti-transfer does not
   lag — it **is** the surface code being read by the role probe (roles ride linear position).
4. **Tied to a reconstruction milestone?** YES. surfX and cross_all_flipped track val_f1
   near-linearly: onset coincides with val_f1 crossing ~0.5 (step 4000), saturation with
   val_f1 ≈ 0.70 (step 26000). Order-code formation is a **byproduct of reconstruction
   improvement** — as the decoder demands word order, the encoder packs it into z via linear
   position, and roles inherit that position-locking.

## Trajectory (agent_patient cells vs training step)

**A_D_s0 (primary, DAE arm; final = certified positive control):**

| step | %bud | val_f1 | ‖z‖ | **surfX** | roleXall | **roleXflip** | roleXlex | winLin | winMlp |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | — | 5.96 | 0.521 | 0.495 | 0.478 | 0.501 | 0.513 | 0.599 |
| 2000 | 5% | 0.408 | 9.67 | 0.519 | 0.495 | 0.480 | 0.499 | 0.504 | 0.591 |
| 4000 | 11% | 0.506 | 9.47 | 0.552 | 0.485 | 0.444 | 0.502 | 0.525 | 0.634 |
| 37360 | 100% | 0.655 | 8.60 | **0.706** | 0.469 | **0.302** | 0.501 | 0.664 | 0.767 |

**A_P90_s0 (tail supplement, paraphrase arm — fills the 11%→100% gap):**

| step | %bud | val_f1 | **surfX** | **roleXflip** | winLin |
|---|---|---|---|---|---|
| 0 | 0 | — | 0.521 | 0.478 | 0.513 |
| 2000 | 6% | 0.470 | 0.534 | 0.463 | 0.514 |
| 4000 | 11% | 0.550 | 0.559 | 0.438 | 0.532 |
| 26000 | 74% | 0.704 | 0.689 | 0.311 | 0.646 |
| 34960 | 100% | 0.709 | 0.692 | 0.309 | 0.649 |

surfX = surface_cross_all_pooled (surface-position code, linear). roleXflip =
role cross_all_flipped (linear; anti-transfer — role probe reads roles BACKWARDS on
parity-flipped families). roleXall = role cross_all_pooled. roleXlex = role
cross_lexical_pooled (stays ≈ 0.50 throughout — the lexical holdout kills the position
ride; consistent with 004/005). winLin/winMlp = within_mean_auc (role readout ceiling).

**Seed & cross-arm consistency:** final surfX = 0.706/0.689 (A_D s0/s1), 0.692/0.702
(A_P90 s0/s1); final roleXflip = 0.302/0.316 (A_D), 0.309/0.298 (A_P90). Both seeds and
both corruption arms show the **same mid-training gradual trajectory** locked to val_f1.
A_P90 (more training pressure) reaches marginally higher surfX/within than A_D — same shape.

## Gates

| Gate | Outcome |
|---|---|
| G1 positive control | **PASS** — A_D_s0 final reproduces `binding_battery_baseline_A_D_s0.json` to 4dp: surfX 0.7063 (t 0.706), roleXflip 0.3021 (0.302), winLin 0.6637 (0.664), z_bag 0.500. Machinery certified deterministic. |
| G2 ckpt identity | **PASS** — loaded step & val_f1 equal inventory for all 18 ckpts (A_D_s0 0/2000/4000/37360; s1 …/6000/37363; A_P90 …/26000/…). |
| G3 z_bag control | **PASS** — z_bag = 0.500 at every ckpt (stimuli valid, lesson 4). |
| G4 norm profile | **PASS** — mean‖z‖ finite, non-degenerate (5.96 untrained → ~9.6 peak at step2000 → ~8.6 final); AUCs are standardized (scale-invariant), so no subtraction norm-lint needed (lesson 3). |

## Predictions → Brier (frozen in PREREG_LITE.md)

| Pred | P | Outcome | Brier |
|---|---|---|---|
| P1 surfX@step2000 (A_D_s0) ≥ 0.62 (order code EARLY) | 0.60 | **FALSE** (0.519 — still at chance at 5% budget) | 0.360 |
| P2 roleXflip@step4000 (A_D_s0) ≤ 0.45 (anti-transfer by mid) | 0.55 | **TRUE** (0.444, marginal) | 0.203 |
| P3 first surfX≥0.62 step ≤ first roleXflip≤0.45 step | 0.70 | **FALSE** (surfX≥0.62 only at final; asymmetric thresholds — they actually move locked) | 0.490 |
| P4 surfX monotone AND largest jump ends ≤ step4000 (front-loaded) | 0.55 | **FALSE** (largest jump is 4000→final; mid-loaded, not front-loaded) | 0.303 |
| P5 G1 pass AND \|surfX(A_D_s1)−surfX(A_D_s0)\|final ≤ 0.06 | 0.85 | **TRUE** (G1 ✓; \|0.689−0.706\|=0.017) | 0.023 |

**Mean Brier = 0.276.** The three misses share one root cause: predictions imported 048's
"atoms crystallize early & front-loaded" prior, but the **order code forms later than atoms**
(mid, ~11-74% budget vs atoms ~77%-done by 11%) and its formation is locked to reconstruction,
not front-loaded. P3's FALSE is a threshold-asymmetry artifact (surfX bar 0.62 vs roleXflip
bar 0.45) — the honest reading of the data is surfX and roleXflip are **simultaneous**, which
is the intended spirit of P3 (surface code does not lag). Direction of the science is clean;
the numeric priors were wrong, as an honest Brier should show.

## What did NOT run / limitations

- **A_D grid is coarse** (milestones at 0/2000/4000/final only — no ms_f1_70; big 11%→100%
  gap). The "gradual mid-training" resolution comes from **A_P90**'s 5-point grid, which is a
  **different corruption arm** (paraphrase, not plain DAE) — flagged. Both arms agree
  qualitatively; a log-spaced-milestone retrain of A_D itself would remove this caveat.
- Single rung (A), d_z=256 **ladder TAE** — NOT SONAR-scale; 2 arms × 2 seeds each.
- Only agent_patient run (genitive/causal/temporal skipped — not needed for the order
  question; the battery's own verdict is INSTRUMENT_FAILURE by construction since genitive
  is weak on this ladder TAE, but the order-code CELLS read directly are the deliverable).
- roleXflip is the sharp anti-transfer metric; cross_lexical stays at chance throughout
  (position ride requires shared vocab) — the phenomenon is a cross_all/shared-vocab effect.
- Box intermediates (emb caches) box-only; 18 batt JSONs + trajectory_073.json + run.log in
  repo `out/`.

## Follow-up worth funding? **Y (modest)**

A cheap retrain of **A_D** (the exact anti-transfer organism) with **log-spaced milestone
ckpts** through 11%→100% (~1 GPU-h) would (a) remove the A_P90 arm caveat and (b) resolve
whether the mid-training rise has any fine abrupt structure hidden in A_D's coarse tail. The
locked surfX↔roleXflip co-movement is a clean handle for a causal test: ablate the linear
surface-position direction at each milestone and check whether role within-ceiling survives
(does role readout depend on the position code as it forms?).
