# RESULT — H2 steering build-up (consolidation / hardening of the 078–080 operator arc) **Tier: T3-exploratory.** Single embedder (SONAR-z, night8 pipeline), LOCAL codex judge (neutral rubric, **0/1320 parse-fail**) as PRIMARY readout, rule-based only as backup. Box tmux `consol_h2` GPU phys0 (CVD=0): 1424 decode jobs, det_ok=True, 434 s. Judge 1320 records, 1242 s. Analysis + codex code-review LOCAL. No promotion above T3. Three deliverables: **(A)** hardened steering on the untested NATURAL-sentence regime; **(B)** a monitored controllable-rewrite tool (`latent_rewrite.py`) with a fabrication flag; **(C)** a local codex code-review of the 079/080 steering+composition code (`CODEX_REVIEW.md`). Frozen predictions in `PREREG_LITE.md`. ## Question 080 found `z + α·v` is causal/specific/invertible with **0% over-steer garble even at α=2** — but on TEMPLATED stimuli, and flagged the break-point as **stimulus-complexity-gated** (natural text should break, cf. 034). H2 fits the IDENTICAL validated operators (negation/tense/vertical) on the same templated TRAIN pairs, then **applies them to held-out NATURAL bases** (entity-rich, 6–26 words, mined from the wnat corpus) over an α grid **extended past 2** ({−1,0,.5,1,1.5,2,3,4,6}) to localize where fluency breaks. Plus: is specificity still perfect on natural text? do operators transfer cross-seed? which are safe-vs-fragile? And a deployable, self-monitoring rewrite primitive on top. ## ★ Headline — on NATURAL text the over-steer break APPEARS and is dose-localized; specificity stays perfect; a round-trip-cosine flag catches the garble The 080 templated 0%-garble result does **NOT** survive on natural sentences. Steering stays fluent and specific through α≈1–1.5, then **breaks into fabrication** (runaway "…not not not not…" / redundant re-marking) at higher α — exactly the stimulus-complexity-gating 080 predicted. The re-encode **round-trip cosine `rt_cos` monotonically tracks the collapse** (0.99→0.96→0.90→0.80→ 0.60→0.29 across α=0→6) and separates codex-fluent from codex-garbled decodes at **AUC 0.845**; nn-cos-to-a-bank does not (AUC 0.78 — it is uninformative on diverse natural text where every sentence is its own nearest neighbour at ~0.25). Specificity is **still perfect (0.00 off-target flip at α=1)** on natural text. Operators are **cross-seed near-identical (cos 0.98–0.995)** on SONAR's fixed basis. The tool refuses **92%** of high-α garble while accepting **100%** of α=1 fluent rewrites. ### Dose–response on NATURAL sentences (codex `fluent` rate | fluent-GATED success | `rt_cos`) | operator | metric | α0 | α0.5 | α1 | α1.5 | α2 | α3 | α4 | α6 | break α | |----------|--------|----|----|----|----|----|----|----|----|---------| | negation | fluent | .90 | .90 | **.90** | .70 | .33 | .10 | .12 | .00 | **2.0** | | negation | success | .00 | .17 | **.70** | .68 | .30 | .05 | .10 | .00 | | | negation | rt_cos | .995 | .984 | .966 | .944 | .906 | .786 | .456 | .236 | | | tense | fluent | .95 | .95 | .88 | .88 | .85 | .60 | .45 | .45 | **3.0** | | tense | success | .00 | .12 | .78 | **.85** | .85 | .60 | .45 | .45 | | | tense | rt_cos | .992 | .987 | .973 | .956 | .933 | .874 | .782 | .348 | | | vertical | fluent | .88 | .92 | .79 | .62 | .58 | .38 | .17 | .04 | **1.5** | | vertical | success | .00 | .00 | **.04** | .04 | .04 | .04 | .00 | .00 | | | vertical | rt_cos | .992 | .982 | .949 | .903 | .857 | .744 | .567 | .299 | | **Over-steer garble is real on natural text** (fluent-rate → 0.00–0.45 at high α; the 080 templated 0% does not hold). The **causal-install window is narrow and dose-localized**: negation installs at α≈1 (0.70) but garbles past α≈1.5; tense is the **most over-steer-robust** operator (fluent 0.85 at α2, breaks only at α3, best success 0.85). The break-point **is stimulus-complexity-gated exactly as 080 predicted** — same operators, natural bases, garble now appears. ### The codex-review fix, caught live (why fluent-gating matters) Ungated negation "success" *rises* to 0.90–0.95 at α1.5–2 — but those are mostly the garbled "…is not …not offering …not providing…" strings whose default polarity read is "negative". The CODEX_REVIEW **BUG #1** fix (gate success on codex-`fluent`) collapses that phantom success to its true window: negation's honest peak is **0.70 at α=1**, and the α≥2 "success" is fabrication, not steering. This is the single most important correction the review produced. ### Fabrication flag — `rt_cos` is diagnostic, nn-cos is not | flag | AUC (fluent vs garbled) | monotone in α | |------|-------------------------|---------------| | **rt_cos** (re-encode round-trip) | **0.845** | ✓ (0.993→0.294) | | nn-cos-to-bank (083/087 metric) | 0.779 | — (flat ~0.2–0.3) | **A finding in itself:** the 083/087 nearest-neighbour-to-a-bank manifold metric **does not transfer to diverse natural text** (natural sentences are semantically unique; max cos to a 600-sentence bank is ~0.25 even for a clean decode). The **round-trip cosine** — decode then re-encode, compare to the steered latent — is the operative fabrication flag: it is ~1.0 for faithful decodes and collapses when steering pushes the latent into non-decodable regions. ### Specificity survives on natural text Off-target attribute-flip rate at α=1 (codex, fluent) = **0.00 for all three operators** (negation doesn't touch tense/number; tense doesn't touch polarity/number; vertical doesn't touch tense). The disentanglement 080 found on templates **holds on entity-rich natural sentences** — the operators are dedicated single-axis directions, not content-entangled. ### Robustness — cross-seed transfer + safe-vs-fragile | operator | cross-seed cos(v₀,v₁) | xseed succ@α1 | break α | verdict | |----------|----------------------|---------------|---------|---------| | negation | **0.984** | 0.70 (= seed0) | 2.0 | safe at α≈1, narrow window; over-steers into "not not not" | | tense | **0.975** | 0.775 (= seed0) | 3.0 | **safest** — most over-steer-robust, strong install (0.85), specific | | vertical | **0.995** | 0.042 (= seed0) | 1.5 | **fragile / does not transfer** to natural text (see below) | **Cross-seed:** operators fit on two disjoint vocab-split seeds are **near-identical (0.975–0.995)** and steer interchangeably (xseed success == seed0 success to 3 decimals). This is the SONAR-specific contrast to **075**: 075's cross-seed non-identifiability (raw cos≈0, needs Procrustes) is a property of a **per-seed-TRAINED** embedder (Ladder-TAE); SONAR is a **fixed** pretrained embedder, so its operators share one basis and transfer **without any rotation alignment**. **Vertical is the honest failure:** it installs the spatial flip at **1.00 on 080's templated "The lamp is above the shelf" frames**, but only **0.04 on natural text** — because natural "above/below" in the corpus is overwhelmingly **non-spatial discourse** ("above all else", "below the poverty line", "see below"), where there is no spatial relation to flip and codex reads `spatial="na"`. The marker operator is real but its *applicability* is frame-bound; it does not generalize to the polysemous natural usages. Documented, and the tool ships it with this caveat. ## (B) The controllable-rewrite tool — `src/latent_rewrite.py` A monitored latent-rewrite primitive: `encode → apply Σ α·v → decode`, with a norm-linter pre-flight (092, cached per operator) and the `rt_cos` fabrication flag that **withholds** off-manifold outputs. Ships only validated operators; **refuses** binding-wall transforms (arg/role swaps) by policy (078/ 079); certifies composition by **decode success**, never additivity cosine (079/CODEX_REVIEW #3). Demo transcript (`out/demo_transcript.txt`, run on box): ``` [ok] negation α1 "The company is expanding into new markets." -> "The company is not expanding into new markets." rt=0.966 accepted [ok] tense α1 "The committee approves the new policy." -> "The committee approved the new policy." rt=0.982 accepted [ok] negation∘tense α1 "The teacher praises the student." -> "The teacher did not praise the student." rt=0.970 accepted ← COMPOSITION [REFUSED] negation α4 "The company is expanding into new markets." raw decode: "The company is not not not not expanding not not not expanding..." rt_cos=0.292 < τ=0.85 -> FABRICATION FLAG, output withheld ← OVER-STEER CAUGHT [ok] vertical α1 "The lamp is above the shelf." -> "The lamp is under the shelf." rt=0.984 [REFUSED] arg_swap -> "binding-wall transform (078/079): not installable; refused." ← BY POLICY ``` The **negation∘tense composition works end-to-end** ("did not praise" = negated + past), the **fabrication flag fires exactly on the over-steered α4 case** (rt_cos 0.29), and the tool refuses what 078/079 proved unreachable. Flag operating point on the held-out sweep: **refuses 91.7%** of high-α (α∈{3,4,6}) codex-garbled decodes, **accepts 100%** of α=1 fluent decodes (τ=0.85, calibrated from this sweep). ## Gates (all PASS) | gate | value | threshold | pass | |------|-------|-----------|------| | ceiling — templated target attr rendered (readout power) | 1.00 | ≥0.85 | ✓ | | ceiling — templated target fluent | 1.00 | — | ✓ | | fabrication floor — templated base has target attr | 0.00 | ≤0.30 | ✓ | | natural α=0 round-trip chrF-to-base | 96.2 | ≥70 | ✓ | | natural α=0 fluent (decoder renders natural text) | 0.913 | ≥0.85 | ✓ | | natural α=0 target-attr rate (fabrication floor) | 0.00 | ≤0.30 | ✓ | | norm-lint (092) per operator | all **PASS** | PASS | ✓ | | codex parse-fail | 0/1320 | ≤5% | ✓ | Readout power is certified by the **positive** ceiling (templated targets decode+read at 1.00, bases at 0.00) — a null here would be instrument failure only if the ceiling failed, and it does not; ppk is unnecessary for a positive of this size. norm-lint flags negation's group-norm mismatch ("center + report cosine") but verdict is PASS; tense/vertical are clean PASS. ## (C) Codex code-review — see CODEX_REVIEW.md Local codex (gpt-5.6-sol, high, read-only, 2.0) | 0.3025 | | P8 | flag refuses ≥0.80 garble AND accepts ≥0.80 α1-fluent | 0.70 | **TRUE** (0.917 / 1.00) | 0.0900 | **Brier = 0.1263** (7/8 directional hits). The lone miss **P7**: I predicted negation would be the most over-steer-robust (extrapolating 080's "negation needs the most drive to *install*"). It is the opposite — negation *installs* hungry but *garbles* early (runaway negation-cue insertion breaks fluency at α2), whereas **tense** is the most robust (redundant past-marking degrades gracefully, break at α3). "Drive needed to install" and "drive tolerated before garble" are different axes; I conflated them. Informative, not a validity failure. ## Verdict (honest) **POSITIVE hardening + a working deployable tool.** 080's clean templated steering **does** break on natural text — the over-steer garble it predicted-but-didn't-see appears here, dose-localized, and is **caught by a round-trip-cosine fabrication flag (AUC 0.845)** while the 083/087 bank metric fails to transfer. Specificity (0.00 off-target) and cross-seed transfer (0.98–0.995, no rotation) survive the natural regime; the operators are dedicated, basis-stable single-axis directions. The payoff is a **monitored latent-rewrite primitive** that composes (negation∘tense), refuses over-steer (92% of garble) and binding-wall swaps, and documents its failure modes — **tense** the safe workhorse, **negation** safe only in a narrow α≈1 window, **vertical** frame-bound and non-transferring to natural polysemous "above/below". The codex review found the steering core correct and one latent scoring bug that this natural-text regime would have tripped — fixed, with its effect shown live. ## Limitations / what did NOT run - **T3 ceiling:** single embedder (SONAR-z), single judge (local codex effort=low), one natural corpus (wnat/nickypro-sonar-sae). Generalization to other decoders/judges untested. - **vertical natural bases are polysemous** (mostly non-spatial "above/below"): its 0.04 natural success measures *frame-transfer failure*, not a regression of the 080 templated 1.00. A curated natural **spatial** bank would separate the two; not built here. - **Random control** is isotropic norm-matched (not covariance/tangent-matched) — a fair but not covariance-aware null (CODEX_REVIEW #6); it garbled ~as much as high-α steer, so it is a weak discriminator on natural text and is not used for a headline claim. - **rt_cos τ=0.85 is calibrated on THIS sweep** (fluent-vs-garbled separation); a deployment would re-calibrate per decoder. AUC 0.845 means the flag is good, not perfect — borderline over-steer (α≈2, rt≈0.90) is the confusable region. - **best-α is test-tuned/optimistic** (CODEX_REVIEW #5); fixed α=1 is the primary reported dose. ## Follow-up worth funding? **Y (modest).** 1. Curated natural **spatial** bank to test whether vertical transfers when the frame is genuinely spatial (isolates frame-transfer from polysemy). 2. Learn a per-sentence **safe-α controller** (stop at the α where rt_cos would cross τ) so the tool auto-doses instead of using a fixed α — the rt_cos curve per sentence already contains the signal. 3. Second embedder/decoder to see if the round-trip-cosine flag's AUC and the tense>negation robustness ordering replicate off SONAR.