# RESULT — Row 058 speech-sonar (EXT) — T3-exploratory

*Run 2026-08-03, tmux c100_058, launched 13:32:26 UTC, DONE 13:47:51 UTC (~15.5 min wall;
smoke passed first, artifacts deleted pre-launch). 0 errors, all 5 stages complete. Box
CVD=0→phys0 hard-guarded. Provenance caveat: the launching manager was killed by a usage
limit AFTER the box run completed but BEFORE harvest; this RESULT.md was written by a
separate harvest agent from out/ artifacts + run.log. The run itself was untouched (DONE
sentinel present, log clean end-to-end); only the manager's in-session context was lost.*

## What ran
piper `en_US-lessac-medium` (deterministic: noise_scale=0, noise_w_scale=0) TTS'd all 3200
unique stimuli_v2 battery sentences (agent_patient 2000 + genitive 1200) + 196 bag words →
3396 wavs @22050 Hz (185 s CPU), resampled to 16 kHz via wave+scipy (torchaudio avoided).
`sonar_speech_encoder_eng` (spenc.eng.pt, 8.0 GB) encoded all 3396 to speech-z (80 s);
`text_sonar_basic_encoder` encoded identical texts to text-z, same run. Cells: (a) alignment
(per-sentence cos + full-3200-pool retrieval both directions), (b) battery UNCHANGED (051
machinery: linear+mlp, seeds 0,1,2, n_boot 1000) on speech_z + text_z measured anchor,
(c) SONAR text decoder on 200 rng(0)-sampled sentences × {text_z, speech_z,
speech_z_centered} (EM strict/loose + chrF), (d) modality offset + centering.

## Gates
- G0 TTS-regime certification: **PASSED pre-freeze** (8 sentences + 17 words: cos mean
  0.928 / min 0.752; 8-way retrieval P@1 = 1.0 both directions; cert_G0_results.json).
- G1 battery internals: PASS — z_bag role = 0.500 exactly in every cell, both reps. Strict
  0.9 gate → formal INSTRUMENT_FAILURE ×2 incl. the text_z anchor (expected and
  pre-registered; 052 A1 measured-anchor logic — the science is speech_z vs text_z deltas).
- G2 decoder validity: PASS — text_z decode chrF 92.9 ≥ 70.
- G3 probe power: PASS — speech_z ap within (mlp) 0.769 ≥ 0.60 → the binding null is a
  real NO, not instrument failure.

## Headline: speech-z IS text-z with a small removable accent — aligned, same
## chance-binding profile (8th NO), decodes through the TEXT decoder at near-text fidelity

**(a) Alignment (full 3200-sentence set)**: cos(speech-z, own-transcript text-z) mean
**0.912** / median 0.927 / p05 0.792 / min 0.339. Retrieval in the full pool (which
contains role-swap minimal pairs with identical bags): speech→text **P@1 = 0.967**,
P@5 = 0.997; text→speech P@1 = 0.858, P@5 = 0.985. The feared failure mode — speech-z
blurring role-swap twins — did not materialize: 96.7% of spoken sentences retrieve their
exact transcript out of 3200 candidates.

**(b) Battery — speech_z vs text_z measured anchor (same sentences, same machinery, same
run)**:

| task | readout | speech within | text within | speech primary (cross+lex-holdout) | text primary |
|---|---|---|---|---|---|
| agent_patient | linear | 0.715 | 0.759 | 0.506 [0.491,0.520] | 0.508 [0.489,0.527] |
| agent_patient | mlp | 0.769 | 0.796 | 0.504 [0.477,0.526] | 0.497 [0.472,0.519] |
| genitive | linear | 0.718 | 0.826 | 0.492 [0.465,0.520] | 0.500 [0.472,0.531] |
| genitive | mlp | 0.771 | 0.835 | 0.482 [0.436,0.529] | 0.493 [0.429,0.543] |

Same regime, point for point: within-construction power present but sub-ceiling in BOTH
reps (speech runs 0.03–0.06 below its text anchor — a mild acoustic tax, not a regime
change); cross-construction + lexical-holdout at chance in BOTH (max CI-lo across all
speech cells 0.491) — the **8th consecutive NO** of the 051–058 arc, and the first
showing the null is modality-general: speaking the sentence changes nothing about the
absence of abstract role binding. z_bag 0.500 everywhere (spoken word bags carry no role
signal either). Curiosity (exploratory): genitive surface-cross is below-chance 0.354 for
text_z but 0.505 for speech_z — the text encoder's surface anti-correlation is not
inherited by the speech path.

**(c) Cross-modal decode (speech-z through the TEXT decoder, n=200)**:

| rep | EM strict | EM loose | chrF corpus | chrF median |
|---|---|---|---|---|
| text_z | 0.670 | 0.670 | 92.9 | 100.0 |
| speech_z | 0.290 | 0.590 | 89.7 | 94.0 |
| speech_z_centered | 0.570 | 0.625 | 92.1 | 100.0 |

The text decoder reads speech-z nearly natively: chrF gap only **3.2** (predicted ≥10 —
P4 FAILED because decode was too GOOD). Qualitative error read (142/200 speech_z
non-exact): dominated by (i) register loss — lowercase/punctuation-stripped output
("no one discussed may's praise of diego"), which alone accounts for the strict/loose EM
gap (0.29→0.59); (ii) ASR-like homophone/near-phone substitutions: Mei→may, Zara→Tsar,
aunt→ant, accountant→counter, supervisor→superintendent; (iii) benign paraphrase shared
with text_z (lawyer→attorney). **No fabrication of new entities or events** — the 022
taxonomy's confabulation mode is absent; errors are perceptual, not generative.

**(d) Modality offset**: ‖μ_sp − μ_tx‖ / mean‖z_tx‖ = **0.149** full-set (cert-set 0.243);
speech-z norms sit lower (mean 0.199 vs 0.221); cos(μ_sp, μ_tx) = 0.964. Centering
(add μ_tx − μ_sp) behaves exactly like the 046/062 language-offset analogue: cos 0.912 →
0.925, retrieval P@1 0.967 → 0.970, and the big effect is on decode surface register —
strict EM **0.29 → 0.57**, chrF 89.7 → 92.1 (median back to 100). The offset largely
encodes "spokenness" (casing/punctuation register), and it is one vector.

## Predictions (frozen) → outcomes, Brier
| pred | p | outcome | Brier |
|---|---|---|---|
| P1 full-set cos mean ≥ 0.85 | .80 | TRUE (0.912) | 0.040 |
| P1b sp→tx retrieval P@1 ≥ 0.80 | .60 | TRUE (0.967) | 0.160 |
| P2 NO binding: speech ap primary CI-lo ≤ 0.55 | .85 | TRUE (max CI-lo 0.491; G3 valid) | 0.023 |
| P3 speech ap within (mlp) ≥ 0.70 | .65 | TRUE (0.769; also within 0.10 of anchor 0.796) | 0.123 |
| P4 speech chrF ≥ 55 AND ≤ text chrF − 10 | .70 | FALSE (89.7 > 82.9 — gap only 3.2) | 0.490 |
| P5 centering: chrF +2 | .60 | TRUE (+2.45; and EM +0.28) | 0.160 |

**Mean Brier = 0.166** (miss = P4, in the good direction: cross-modal transfer is far
tighter than predicted).

## Honest verdict
YES on the row's question, with the arc-consistent null intact: SONAR's speech encoder
drops spoken sentences almost exactly where their transcripts live (0.91 cos, 97% exact
retrieval against role-swap twins), carries the same sub-ceiling-power/chance-binding
battery profile as text-z (8th NO, now modality-general), and is read by the TEXT decoder
at near-text fidelity with ASR-flavored — never confabulatory — errors. The modality gap
is well-approximated by a single mean offset whose removal doubles strict exact-match by
restoring written register. T3-exploratory.

## Limitations / what did not run
- One deterministic neural TTS voice (piper lessac-medium), not natural speech: no speaker
  variation, prosody variation, disfluency, or noise. G0 certifies the battery is
  in-regime for THIS voice; all numbers are upper bounds on natural-speech behavior.
- Killed-manager provenance (see header): harvest from artifacts only; nothing anomalous
  found, but no in-session eyes were on stages 3–5 as they ran.
- Loose EM only strips case + trailing period; some "errors" counted at strict EM are
  pure punctuation. chrF is the register-robust number.
- 022 fabrication taxonomy applied qualitatively to 200 samples, not re-run as a classifier.
- Strict battery gate is formal INSTRUMENT_FAILURE for both reps (as pre-registered);
  binding conclusions rest on the measured text_z anchor comparison, per 052 A1.

## Follow-up worth funding? N (as a row)
The cross-modal geometry result is clean and the binding arc doesn't need a 9th cell.
Fold-ins if ever revisited: natural speech (LibriSpeech/CommonVoice) to stress the
TTS-regime caveat; non-English speech encoders vs the 046/062 language offsets (is the
"modality offset" one axis or per-language?); offset-as-register probe (does μ_sp − μ_tx
decode as lowercase style on text-z too?). The fundable question remains the arc-level
synthesis: what training signal WOULD install binding.

## Artifacts
Repo `out/`: results_058.json, cert_G0_results.json, decode_samples_058.json,
binding_battery_{speech_z,text_z}.json, BATTERY_RESULTS_058.md, z_speech.npz + z_text.npz
(13 MB each), z_index.json, run.log. Box: audio/ (~250 MB wavs) + venv_tts + voices
retained on box only; fairseq2 cache +8.0 GB (spenc.eng.pt).
