# RESULT — 080 sonar-causal-steering (Campaign100 MANIFEST row 080, BLOCK J)

**Tier: T3-exploratory.** Single embedder (SONAR z, night8 pipeline), templated minimal-pair
stimuli, **rule-based deterministic attribute readout** (controlled vocab, synonym-aware) as
PRIMARY. Box tmux `c100_080` GPU phys0 (CVD=0). Decode = 5920 jobs, det_ok=True, **86.4 s**.
Analysis local. Codex fluency cross-check = OPTIONAL follow-up (not run; validity gate already met
by ceiling + baseline + chrF, see Gates). No promotion above T3.

## Question
The preregistered CAUSAL steering test: take a validated linear operator `v` (fit on disjoint TRAIN
pairs), add `α·v` to a held-out base z, decode, measure the DOSE–RESPONSE of the target attribute
vs α ∈ {−1,−0.5,0,0.5,1,1.5,2}. Does steering causally & monotonically install the attribute? Where
does it saturate / break off-manifold (cf. 034 α=2 garble, 038 antipode void)? Is it SPECIFIC (flips
only the target, not the other attributes)? INVERTIBLE (−v undoes)? Is a RANDOM norm-matched push
inert (certifying v is special)?

## Operators (4 validated) + families
Grammatical SVO "The {subj} {verb} the {obj}." (attrs polarity/tense/number): **negation**
(aff→neg), **tense** (pres→past), **number** (sing→plur). Spatial copular "The {a} is above the {b}."
(attrs vertical/tense): **vertical** (above→below). `v_T = mean(z_target−z_base)` over 64 TRAIN
pairs; TRAIN vocab disjoint from TEST. 90 forward + 50 reverse held-out bases per operator.

## ★ Headline — steering is CAUSAL, MONOTONE, SATURATING, PERFECTLY SPECIFIC, INVERTIBLE, and v is SPECIAL
Adding `α·v` **causally installs the target attribute** with a clean threshold-then-plateau dose
curve; a **norm-matched random direction does nothing** (0.00 at every α); steering the target
attribute produces **zero collateral flips** of the other attributes; and **−v perfectly undoes**
(reverse bases recover the +pole at α=−1). On these templated stimuli there is **NO over-steer
garble even at α=2** — the single prereg miss (P2), and itself an informative negative.

### Forward dose–response — success (target-attribute installed) vs α
| operator | −1 | −0.5 | 0 | 0.5 | 1.0 | 1.5 | 2.0 | garble(any α) |
|----------|----|----|----|----|----|----|----|----|
| negation | 0.00 | 0.00 | 0.00 | 0.00 | 0.93 | 0.94 | **0.96** | 0.00 |
| tense    | 0.00 | 0.00 | 0.00 | 0.48 | **1.00** | 1.00 | 1.00 | 0.00 |
| number   | 0.00 | 0.00 | 0.00 | 0.19 | **1.00** | 1.00 | 1.00 | 0.00 |
| vertical | 0.00 | 0.00 | 0.00 | 0.76 | **1.00** | 1.00 | 1.00 | 0.00 |

Monotone non-decreasing 0→peak for all four (mean Δ(α1−α0) = **+0.98**). Sharp sigmoid: essentially
0 below α=0.5, saturated by α=1 (negation needs the most drive — still climbing at α=2, 0.96 —
consistent with 034's under-scaled natural-negation offset). α=0 identity = **0.00** everywhere
(baseline fabrication floor is ZERO) and decodes the original (chrF-to-base **95.4**).

### cos(steered z, true target z) vs α — the "right dose" is α≈1
Peaks at α≈1 for all ops (negation 0.973 / tense 0.985 / number 0.982 / vertical 0.995), falling on
both sides — the steered latent is geometrically closest to the true attribute-flipped point at α=1,
independently corroborating the decode readout.

### Specificity @α=1 — flip the target ONLY
| operator | target success | off-target flip rate | other attrs measured |
|----------|---------------|----------------------|----------------------|
| negation | 0.93 | **0.00** | tense, number |
| tense    | 1.00 | **0.00** | polarity, number |
| number   | 1.00 | **0.00** | polarity, tense |
| vertical | 1.00 | **0.00** | tense |

Zero collateral flips — steering negation changes polarity and nothing else; 036/038's off-manifold
collateral-flip worry does **not** materialize at α=1 on these operators.

### Invertibility (reverse bases, −v undoes toward +pole)
undo-success at **α=−1 = 1.00 for all four operators** (α=−0.5 partial: negation 0.94 / number 0.72 /
tense 0.62 / vertical 0.38). +α on a −pole base = 0.00 (correctly does not over-install). +v and −v
are clean inverses, replicating 078's marker-invertibility and extending it to negation/tense/number.

### Random-direction control — v is special
Norm-matched Gaussian `v_rand` (‖v_rand‖=‖v‖, cos(v_rand,v)≈0): target-attribute success **0.00 at
every α ∈ {0.5,1,1.5,2} for all four operators**. The attribute install is not a generic
large-norm-push artifact; it is specific to the fitted operator direction.

### Manifold departure (nn-cos of steered z to a 900-sentence real templated bank)
| operator | α=0 | α=1 | α=2 |
|----------|-----|-----|-----|
| negation | 0.926 | 0.909 | 0.863 |
| tense    | 0.913 | 0.911 | 0.879 |
| number   | 0.936 | 0.929 | 0.893 |
| vertical | 0.976 | 0.967 | 0.905 |

nn-cos peaks at α≈0 and falls with |α| (mean α=0 0.938 → α=2 0.885, Δ≈0.053). Steering does drift the
latent off the real manifold at high dose — but on this templated domain the drift is mild enough
that decode stays **fluent and correct** (0% garble). Contrast 034 (natural sentences garbled at
α=2) and 038 (antipode void): the break-point is stimulus-complexity-dependent, not reached here.

## Offset geometry (norm profile, lesson 3)
| operator | ‖v‖ | ‖z_base‖ | diff_align (linearity) | cos(v_rand, v) |
|----------|-----|----------|------------------------|----------------|
| negation | 0.059 | 0.184 | 0.847 | −0.01 |
| tense    | 0.046 | 0.186 | 0.838 | +0.01 |
| number   | 0.045 | 0.184 | 0.816 | +0.02 |
| vertical | 0.062 | 0.194 | 0.962 | −0.04 |

‖v‖ is ~24–33% of ‖z‖; α=1 adds one operator-norm. diff_align 0.82–0.96 confirms these are coherent
linear directions (vertical the tightest, as in 078).

## Gates (all PASS)
| gate | value | threshold | pass |
|------|-------|-----------|------|
| ceiling success (decode true target) | 0.998 (0.993–1.00) | ≥ 0.85 | ✓ |
| α=0 identity decode chrF-to-base | 95.4 | ≥ 80 | ✓ |
| baseline (α=0) target-attr rate (fabrication floor) | 0.00 | ≤ 0.30 | ✓ |
| det check | True | — | ✓ |
| random-direction null @α=1 | 0.00 | ≤ 0.15 | ✓ |
Ceiling ~1.0 + baseline 0.00 together are the decode-validity / fabrication check (lesson 4): the
success signal is real applied structure, not decoder confabulation. Readout is synonym-aware
(SONAR decodes "below"→"under", "above"→"over"; the extractor matches the semantic class — caught in
smoke, where a literal-word extractor spuriously scored ceiling 0.33 before the fix).

## Verdict (honest)
**POSITIVE and clean.** For all four validated operators, `z + α·v` is a **causal, dose-responsive,
attribute-specific, invertible** steering control, and a random push of equal norm is inert. This is
the causal complement to 033/034/078's correlational "linear offset operator" finding: the
directions don't just *describe* the attribute geometry, they *install* it. The dose curve is a
sharp threshold (≈0 below α=0.5) to a plateau (saturated by α=1; negation lags, peaking at α=2),
with the latent geometrically closest to the true target at α≈1. **The prereg's one miss is
scientific signal, not failure:** predicted over-steer garble at α=2 (extrapolating 034) did NOT
occur — on clean templated stimuli the operators steer without breaking fluency across the whole
tested range; the off-manifold break-point seen in 034 (natural text) / 038 (antipode) is
stimulus-complexity-gated and beyond α=2 here. Specificity (0.00 off-target) and the random null
(0.00) are the strongest cells: they certify these are *disentangled, dedicated* causal directions.

## Brier (frozen priors, PREREG_LITE, exact)
| pred | statement | prior P | outcome | (P−o)² |
|------|-----------|---------|---------|--------|
| P1 | monotone & Δ(α1−α0)≥0.4 | 0.75 | **TRUE** (Δ=0.98, monotone) | 0.0625 |
| P2 | over-steer break α=2 & garble↑ | 0.70 | **FALSE** (no break; 0% garble at α=2) | 0.4900 |
| P3 | off-target flip ≤0.20 @α=1 | 0.65 | **TRUE** (0.00) | 0.1225 |
| P4 | random success ≤0.15 @α=1 | 0.85 | **TRUE** (0.00) | 0.0225 |
| P5 | nn-cos(α2) ≤ nn-cos(α0)−0.05 | 0.70 | **TRUE** (0.885 ≤ 0.938−0.05) | 0.0900 |
| P6 | undo-success(α=−1) ≥0.40 | 0.60 | **TRUE** (1.00) | 0.1600 |

**Brier = 0.1579** (5/6 directional hits). The loss is dominated by P2: I over-anticipated 034's
natural-text over-steer garble on templated stimuli, where steering is robust to α=2.

## Limitations / what did NOT run
- Templated stimuli (controlled vocab) — the 0% garble / perfect specificity likely **overstate**
  robustness vs natural text (034 shows garble emerges there); the clean numbers are a best-case
  ceiling for these operators. Natural-sentence dose-response is the obvious follow-up.
- Rule-based readout (deterministic, synonym-aware) is PRIMARY; a LOCAL codex fluency + attribute
  cross-check (judge_steer.py is written and ready) was **not run** — the validity gate is already
  met by ceiling 0.998 + baseline 0.00 + chrF 95.4 + 0% rule-based garble, so codex is
  belt-and-suspenders, deferred as optional.
- α grid capped at |2|; the true off-manifold break-point (where garble finally sets in) is beyond
  the tested range for templated stimuli — not localized here.
- 4 operators, all previously validated as linear (033/034/078); this tests CAUSALITY of the
  already-linear set, not whether steering can install a NON-linear transform (078 says it can't:
  arg-swaps hit the binding wall — untested here by design).

## Follow-up worth funding? **Y** — natural-sentence dose-response (where the over-steer break
should reappear, per 034) + push α past 2 to localize the fluency break-point + steer a
CLOSED-CLASS marker on natural text. The causal claim is now clean at T3; a natural-text replication
with the codex fluency gate would be the promotion path above T3.
