# RESULT — 023 offmanifold-dose

**Tier: T3-exploratory.** Box decode (400 sentences × 19 conditions = 7600 greedy decodes +
re-encode) + LOCAL codex 3-way judging (1900 V1 labels, **0 parse failures**) + analysis, all
run per the FROZEN `PREREG_LITE.md`. No prereg deviations.

## Verdict — NEGATIVE on the headline hypothesis (+ informative direction-dependent shape)

Pushing a real sentence's SONAR embedding off-manifold by a controlled dose does **NOT** open a
"fail-open band" of fluent-but-wrong reconstructions for **any** of the three perturbation
directions. The pre-registered band (altered-but-fluent > 0.30 while garbled < 0.20) **never
occurs** — at every (direction, dose) where fabrication rises, garble rises with it. All four
frozen predictions resolve FALSE; **Brier = 0.3881** (mis-calibrated — we were confident, at
p(a)=0.80, in a band that does not exist). The instrument is valid: all sanity gates pass
(G1 dose-0 ≈ 021 baseline; G2/G3 cosines monotone; G4 hand-vs-codex 95%).

The real yield is the **dose-response map**: the decoder is robust to off-manifold perturbation
up to moderate doses, and its failure MODE is strongly **direction-dependent** — random
directions are near-inert, off-manifold-normal directions break into **garble**, and only the
chord toward another real sentence produces fabrication, and even then it arrives fused with
garble rather than as a clean fail-open zone.

## Dose-response table (codex V1, n=100 paired ids per condition)

`faith/alter/garb` = 3-way label fractions; `mlp` = mean chosen-token logprob; `chrf` vs orig;
`cosIN` = cos(z, z'); `cosRE` = cos(z, z_rec); `degen` = degenerate-flag rate.

**dose 0 (shared baseline):** faith 0.95 · alter 0.01 · garb 0.04 · mlp −0.143 · chrf 96.3 · cosRE 0.996 · degen 0.00

| dir | dose | faith | alter | garb | mlp | chrf | cosIN | cosRE | degen |
|-----|------|-------|-------|------|-----|------|-------|-------|-------|
| **rand** | 0.05 | 0.97 | 0.01 | 0.02 | −0.14 | 96.0 | 0.999 | 0.995 | 0.00 |
| | 0.10 | 0.98 | 0.02 | 0.00 | −0.15 | 96.1 | 0.995 | 0.996 | 0.00 |
| | 0.20 | 0.99 | 0.01 | 0.00 | −0.15 | 95.8 | 0.981 | 0.995 | 0.00 |
| | 0.35 | 0.94 | 0.05 | 0.01 | −0.17 | 94.6 | 0.944 | 0.990 | 0.00 |
| | 0.50 | 0.87 | 0.07 | 0.06 | −0.20 | 93.2 | 0.895 | 0.988 | 0.00 |
| | 0.75 | 0.72 | 0.20 | 0.08 | −0.30 | 87.6 | 0.801 | 0.961 | 0.00 |
| **chord** | 0.05 | 0.97 | 0.03 | 0.00 | −0.15 | 96.0 | 0.999 | 0.995 | 0.00 |
| | 0.10 | 0.97 | 0.02 | 0.01 | −0.15 | 95.4 | 0.997 | 0.994 | 0.00 |
| | 0.20 | 0.95 | 0.05 | 0.00 | −0.16 | 94.7 | 0.986 | 0.993 | 0.00 |
| | 0.35 | 0.86 | 0.10 | 0.04 | −0.21 | 91.4 | 0.948 | 0.984 | 0.00 |
| | 0.50 | **0.42** | **0.22** | **0.36** | −0.40 | 76.1 | 0.876 | 0.911 | 0.00 |
| | 0.75 | **0.00** | **0.60** | **0.40** | −0.72 | 24.5 | 0.669 | 0.268 | 0.02 |
| **offman** | 0.05 | 0.96 | 0.01 | 0.03 | −0.14 | 96.0 | 0.999 | 0.996 | 0.00 |
| | 0.10 | 0.96 | 0.02 | 0.02 | −0.14 | 95.9 | 0.997 | 0.995 | 0.00 |
| | 0.20 | 0.92 | 0.03 | 0.05 | −0.15 | 95.6 | 0.989 | 0.995 | 0.00 |
| | 0.35 | 0.86 | 0.06 | 0.08 | −0.16 | 95.1 | 0.970 | 0.991 | 0.00 |
| | 0.50 | 0.76 | 0.09 | 0.15 | −0.19 | 94.3 | 0.948 | 0.982 | 0.00 |
| | 0.75 | 0.57 | 0.08 | **0.35** | −0.26 | 89.9 | 0.910 | 0.955 | 0.00 |

## Band characterization — the fail-open band does not exist

Registered band = {dose : altered > 0.30 AND garbled < 0.20}. **Empty for all three directions:**

- **rand** — near-inert. Altered peaks at only **0.20** (dose 0.75, cosIN already 0.80); garble
  never exceeds 0.08. Random 1024-d directions are ~orthogonal to the decodable subspace, so most
  of ‖δ‖ is wasted; degradation is graceful and dominated by faithful loss, not fabrication.
- **chord** (toward another real z) — the **only** potent direction, and the sharpest failure:
  faithful 0.86 → 0.42 → 0.00 across doses 0.35/0.50/0.75. But fabrication and garble rise
  **together** — at dose 0.50 altered 0.22 / garbled 0.36; at 0.75 altered 0.60 / garbled 0.40.
  Never fluent-wrong-dominant with garble suppressed → no band. cosRE collapses to 0.27 at 0.75
  (total loss of the original). This is a *mush zone*, not a fail-open band.
- **offman** (SAE residual normal) — preferentially breaks into **garble** (0.35 at dose 0.75)
  with almost no fabrication (altered ≤ 0.09 everywhere). Pushing along the off-manifold normal
  degrades fluency rather than inventing plausible content.

Even the softer view — altered / (faithful+altered), the share of *fluent* outputs that are wrong
— only reaches 0.34 (chord 0.50) and 1.00 (chord 0.75, where almost nothing stays fluent); rand
tops out at 0.22 and offman at 0.12. Fabrication is scarce off-manifold except along the chord.

## Predictions (frozen) → outcomes → Brier

| id | statement | p | outcome |
|----|-----------|---|---------|
| (a) | fail-open band exists (altered>0.30 & garbled<0.20) for ≥1 direction | 0.80 | **FALSE** — empty for all 3 |
| (b) | chord band wider than rand | 0.60 | **FALSE** — both bands empty |
| (c) | rand shows a cliff (garbled>0.50 within 2 doses of altered first >0.30) | 0.50 | **FALSE** — rand altered never >0.30 |
| (d) | altered outputs' mean re-encode cosine > 0.90 (cosine-invisible fail-open) | 0.55 | **FALSE** — mean 0.678 (N=168) |

**Brier = 0.3881** (mean squared error over a–d). Worse than an all-0.5 reference (0.25): the prior
over-weighted a "wide cosine-invisible fabrication zone" that the data refute. The confident miss
on (a) (0.80) and (b) (0.60) dominate the loss.

**Nuance on (d) — refuted globally, but direction-split is informative.** The global altered
re-encode cosine (0.678) is dragged down by chord, which supplies 102 of 168 altered outputs at a
mean cosRE of just **0.52** (chord fabrications ARE cosine-visible — a cos gate would catch them).
The *few* rand (n=36, mean 0.91, 67% > 0.90) and offman (n=29, mean 0.93, 83% > 0.90) fabrications
are cosine-near-invisible — consistent with 021/022's "gate-blind fabrication" but they are rare.
So off-manifold, cosine-invisible fabrication exists only as a thin tail, not a regime.

## Gates / sanity — ALL PASS (valid instrument, NOT instrument failure)

- **G1 dose-0 sanity: PASS.** altered 0.01 < 0.10, garbled 0.04 < 0.05. Off-clean = 5% ≈ 021's
  3.6% semantic-fab baseline (021 re-decoded a different corpus with beam+nucleus; the ballpark
  match confirms rubric+pipeline are not manufacturing fabrication).
- **G2 cosIN monotone-decreasing in dose:** PASS all 3 directions (arithmetic check on the
  perturbation, as expected).
- **G3 re-encode cosine non-increasing in dose (within 0.02 tol):** PASS all 3 directions
  (semantic fidelity degrades monotonically with dose).
- **G4 hand-vs-codex:** manager hand-labeled 20 stratified pairs BEFORE reading codex →
  **19/20 = 95% agreement.** Single disagreement `offman:0.5:1108` ("…phonology, syntax, syntax,
  and semantics.") hand=faithful (meaning intact, minor reduplication) vs codex=garbled
  (repetitive output) — a genuine altered/garbled boundary call on a doubled token.
- **Judge reliability (V1 vs differently-worded V2, n=150):** raw agreement 0.827, Cohen's
  κ = 0.581 (moderate). Disagreement concentrates on the altered-vs-garbled surface boundary
  (V2 was more lenient, calling repetition "ok/wrong-but-clean" not "broken") — the same boundary
  as the G4 miss. The faithful-vs-not split is much more stable.

## Caveats / limitations

- **Greedy-only** decode (no beam/nucleus); a sampling decoder could populate the fluent-wrong
  region differently. This is the 020/022-validated path, chosen for determinism.
- **One corpus** (021's cpool = nickypro/sonar-sae C4-real family), one SAE checkpoint
  (cpool h16384 k32 BatchTopK, FVU 0.44) defining the offman direction. Off-manifold-normal
  findings are conditional on that SAE's residual.
- **Judge subsample = 100 ids × 19 conditions** (paired). Condition fractions have ~±0.05 binomial
  SE at n=100; small cells (e.g. rand altered=0.05) are noisy — the qualitative shape (robust →
  chord-mush / offman-garble) is what is load-bearing, not any single 2nd-decimal.
- Single LLM judge (codex, reasoning=low). κ=0.58 rubric sensitivity means the altered/garbled
  split carries real annotator noise; faithful mass (the bulk) is reliable.
- T3-exploratory: a shape/negative result, not a promotable claim.

## Follow-up worth funding? — **Weak Y**

The clean finding "the decoder does not fail open off-manifold; it fails *closed* into garble
except along inter-sentence chords" is a useful safety-relevant map. A worthwhile extension is a
**sampling-decoder** rerun of the chord direction only (where fabrication concentrates) at fine
dose spacing around 0.35–0.6, to see whether sampling widens the fluent-wrong shoulder that greedy
suppresses. Lower priority than the 024/086 prior-implausibility detector thread.
