# 018 function-word-carriers — RESULT

Tier: **T3-exploratory**. CPU-only, binding_death states READ-ONLY. Ran 2026-08-02, box
`campaign100/018-function-words/`, tmux `c100_018`. Chain wall ~50 s (+6 s span precompute).

## Verdict
**NEGATIVE (certified).** Function-word / grammar-token STATES carry **no** construction-internal
role-or-order code that transfers across vocabulary. Across 8 families × 6 layers × 2 targets, the
`funcword`-span cells sit at the random-content-token noise floor (AUC 0.40–0.69, every CI-lo ≤
0.55, no consistent layer trend), while the moving-filler positional anchor reads the order code at
**AUC 1.000 everywhere** and both planted-signal power certs pass on the funcword state itself. So
the grammar tokens `by / was / who / whom / 's / of / has` do **not** retain, in their linear state,
which noun is the agent — even within a single construction with only a lexical (vocab) holdout.
This complements 012 (role-consistent attention *reads* these cues) + 014 (role rides the filler
tokens): the cue TOKEN is a conduit, not a store.

## What ran
- **Probe** = battery linear logistic (`binding_battery._fit/_pred`, TRAIN-only standardize) on the
  per-item MEAN state of a position class. Reused `binding_battery.{_fit,_pred,_auc,
  prop_bootstrap_auc,_env}` as a LIBRARY (zero shared-file edits). numpy-only, RAW NUMBERS.
- **Tasks/families (8)**: agent_patient {active,passive,cleft,objrel,nominal} + genitive
  {s_genitive,of_genitive,have}, each probed SEPARATELY (construction-internal).
- **Cell** = WITHIN-family LEXICAL holdout: train `family & split==train` (300 items, 150:150),
  test zero-shot `family & split==test` (100 items, 50:50, 50 test props, all count 2). Prop-cluster
  bootstrap n_boot=1000, 3 seeds, `_env` envelope. tick-17 BLOCK RULE applied (VERIFIED: funcwords
  structural; only active proper-name items lack the determiner span — dropped per prop, see below).
- **Position classes**: `funcword` (family grammar tokens, SPM-piece matched; possessive `'`+`s` by
  adjacency), `fillerA` / `fillerB` (each entity's own span — positional anchors), `randctl`
  (count-matched random content tokens ∉ fillers∪funcword, seeded — noise floor), `allpool`.
- **Layers**: {L4,L8,L12,L16,L20,L24n}. **Targets**: `role`=`_y_role` (focal=alpha-first filler);
  `order`=raw `A_surface_first`.

## Gates
- **G-align** PASS — re-tokenized T==meta.T (0 mismatch, 3200 items) and A/B pieces==meta (0
  mismatch); funcword span nonempty for every item EXCEPT 40 active proper-name items ("Mei sued
  Marco.", no determiner) — handled by the block rule (see incidents).
- **G-power** PASS (both cells) — planted construction-invariant role signal INTO the funcword state,
  rerun full within-family pipeline: `agent_patient/passive/L12` additive certified @ε=0.15
  (recovery 0.906); `agent_patient/nominal/L24n` norm-preserving rotation certified @θ=0.6
  (recovery 0.968). The funcword null is therefore powered, not blind.
- **G-anchor** PASS (decisive) — `fillerA|order` = `fillerB|order` = **1.000 at every layer/family**
  (min 1.000). The instrument reads the positional order code perfectly from the token that
  physically occupies the position. (This is expected/near-trivial — A_surface_first is defined by
  A's position and A's token carries it — but it is the required "known code shows" control; the
  non-trivial power control is the G-power plant on the funcword state.)

## Numbers

**funcword | role (AUC per layer; noise floor = randctl max per family in brackets):**

| family | L4 | L8 | L12 | L16 | L20 | L24n | [randctl max] |
|---|---|---|---|---|---|---|---|
| active | 0.447 | 0.454 | 0.631 | 0.520 | 0.504 | **0.688** | [0.563] |
| passive | 0.548 | 0.535 | 0.530 | **0.616** | 0.570 | 0.507 | [0.529] |
| cleft | 0.564 | 0.475 | 0.473 | 0.478 | 0.465 | 0.476 | [0.592] |
| objrel | 0.459 | 0.493 | 0.410 | 0.399 | 0.490 | 0.490 | [0.552] |
| nominal | 0.505 | 0.484 | 0.537 | 0.560 | **0.617** | 0.572 | [0.522] |
| s_genitive | 0.478 | **0.608** | 0.477 | 0.500 | 0.456 | 0.512 | [0.501] |
| of_genitive | 0.550 | 0.506 | 0.480 | 0.464 | 0.459 | 0.458 | [0.480] |
| have | 0.510 | 0.580 | 0.473 | 0.434 | 0.587 | 0.514 | [0.546] |

- Global `funcword|role` argmax = **active/L24n 0.688** [0.545, 0.830] — a determiner cell on the
  reduced (proper-name-dropped) set, CI spanning chance. Passive `by` peaks at 0.616 [0.55, 0.69].
- Global `funcword|order` argmax = of_genitive/L16 0.618 [0.47, 0.76]; AP funcword|order all ≤0.61.
- `allpool` max: role 0.561, order 0.646 — the whole-sentence pooled readout is ALSO ~chance under
  this discipline (secondary finding, below).
- `fillerA`/`fillerB` |order = 1.000 (min) — the positional anchor dominates every funcword cell,
  every layer, both targets.

**Read:** funcword peaks (active 0.688, nominal 0.617, passive 0.616) are within the multiple-
comparison noise of the randctl floor (per-family randctl role_max 0.48–0.59 over the same 6-layer
search); no CI-lo clears 0.55, no monotone layer profile. There is at most a *whisper* at passive
`by` / nominal `'s`/prep, but nothing survives as a code.

## Predictions & Brier (frozen in PREREG_LITE)
| # | prediction | P | outcome | Brier |
|---|---|---|---|---|
| a | ANY funcword|role > 0.7 in ≥1 family/layer | 0.15 | **FALSE** (max 0.688) | 0.0225 |
| b | ANY funcword|order (A_surface_first) > 0.9 | 0.08 | **FALSE** (max 0.618) | 0.0064 |
| c | filler anchor > EVERY funcword cell, everywhere | 0.85 | **TRUE** (fillerA=1.000 > all) | 0.0225 |
| d | passive `by` is the strongest funcword|role cell | 0.30 | **FALSE** (argmax=active determiner) | 0.0900 |
| | | | **mean** | **0.0354** |

Well-calibrated (the arc's strong state-code-null prior paid off). The (d) miss is benign: the
argmax landed on active's degenerate determiner cell by noise (wide CI); passive `by` = 0.616 was
the strongest *role-marking* funcword, but still short of a code.

## Secondary yield (disclosed design clarification — see incidents)
Within a single family, `_y_role` (focal, alpha-first) and surface order are a deterministic
bijection, so the role/order *dissociation* is realized as **focal-referenced (needs the two nouns'
alphabetical rank — a lexical fact) vs A-referenced position**. Under LEXICAL holdout, the
focal-referenced target is **not recoverable from ANY span** (funcword, fillerA, fillerB, allpool all
≈ 0.5) because SONAR states of held-out vocab don't linearly encode alphabetical rank; the ONLY
survivor is the raw positional order code, readable at 1.000 but ONLY from the token that occupies
the position (fillerA/B), not from anything pooled or fixed-position. This sharpens the arc's
"within-construction role is a lexical/positional lookup, not an abstract variable": strip the
vocabulary and the lookup collapses to bare token position.

## Limitations
- Linear mean-state probe only (matches the battery/arc regime); a nonlinear or attention-pooled
  read of the funcword state was not tried here (016 exhausted nonlinear reads on filler/pooled
  states → null). The G-power cert licenses the null only against planted signals of the certified
  size (ε=0.15 additive / θ=0.6 rotation).
- The order anchor (fillerA=1.000) is near-trivial (positional); the load-bearing control is the
  funcword-state plant cert.
- active `funcword`="the" is a degenerate determiner cell (no role cue) — kept as an in-experiment
  null; its 40 proper-name items carry no determiner and were dropped by the block rule.
- N per within-family test cell is 100 items / 50 props (CIs ~±0.07–0.13); 6-layer × 8-family search
  inflates the max — read the maxes as an upper noise envelope, not point estimates.

## Follow-up worth funding? **N (low priority).**
Grammar tokens are a conduit for the role-reading attention (012), not a store — consistent with
014's "role rides the fillers." A funcword-state probe adds no transferable role code, and the
nonlinear escape hatch is already closed (016). The only mildly-live thread is the passive-`by` /
nominal-`'s` whisper (~0.6): if anything, test it with the 012 attention-write direction (does a role
head WRITE agent identity into `by`?) rather than a fresh linear state probe — but the prior is low.
