# RESULT — 046 crosslingual-atoms: does the English-trained k=32 SAE fire the same atoms on the same meanings across languages?

**Tier: T3-exploratory.** GPU CVD=0→phys0 (guarded, 35 MiB pre-claim), tmux c100_046,
full chain 58 s (encode 6×2009 = 49.5 s, analysis 6 s), smoke passed first, DONE 05:27:49Z,
SELF-HARVESTED in-session 2026-08-03. Prereg adopted unchanged from the usage-limit-killed
predecessor + pre-compute operationalization amendment (both in PREREG_LITE.md).

**Verdict: ★ STRONG POSITIVE — atom identity is language-invariant on this instrument.**
On FLORES-200 parallel text (2009 rows × eng/deu/jpn/tur/zho/arb, SONAR z), the
activation-correlation matcher (044 lesson: decoder cosine is NOT identifiable; this is the
fix, and here it is certified) maps **every candidate atom to itself in all 5 en↔X pairs,
both directions, raw AND centered: identity rate 1.000** (n_cand 268–302 per pair, top-32
fire ≥20 in both langs), with clean controls (self 1.0, translation-scale-noise 1.0 at
displacement = 51% of ‖z‖, scrambled-alignment 0.0035 ≈ chance 1/289). The dictionary's
frequent atoms are meaning-indexed, not English-surface-indexed. But firing is only partially
shared per sentence: translation pairs share a median 21–33% of their top-32 sets (vs 1.6%
null) — same atom inventory, substantially different per-sentence selections. Mean Brier
0.179 (P1 T, P2 T, P2b F at ceiling, P3 T, P4 F by 0.004).

## Gates (all PASS, certified at full n per 044's smoke-vs-full lesson)
- **G1 ckpt identity**: pile-4k FVU **0.7312** (band 0.731±0.02; ≡045's 0.7312).
- **G2 matcher certification**: self-match 1.000 (n=364); Gaussian noise at translation-scale
  displacement (median ‖z_deu_cen − z_eng‖ = 0.1204 = **0.509×** mean raw ‖z‖ 0.236) → 1.000
  (n=274); scrambled-alignment (derangement) → **0.0035** (n=289). Instrument discriminates.
- **G3 encoder/alignment**: retrieval P@1 en↔X = 1.000/1.000 all pairs except jpn 0.999/0.999
  (min 0.999 ≥ 0.8; above 062's 0.884 floor).
- **G4 norm lint**: mean ‖x_n‖ ratio max/min **1.041** < 5 (raw-space 0.245/0.236 = 1.037).

## (a) Atom identity across languages (matcher, ReLU pre-act profiles over 2009 aligned rows)
| pair | raw | raw rev | centered | n_cand | self-corr med (post-hoc) | runner-up med | margin med/min |
|---|---|---|---|---|---|---|---|
| en↔deu | 1.000 | 1.000 | 1.000 | 289 | 0.872 | 0.193 | 0.662 / 0.136 |
| en↔jpn | 1.000 | 1.000 | 1.000 | 268 | 0.791 | 0.189 | 0.586 / 0.256 |
| en↔tur | 1.000 | 1.000 | 1.000 | 277 | 0.842 | 0.192 | 0.626 / 0.115 |
| en↔zho | 1.000 | 1.000 | 1.000 | 271 | 0.778 | 0.193 | 0.570 / 0.268 |
| en↔arb | 1.000 | 1.000 | 1.000 | 276 | 0.852 | 0.194 | 0.646 / 0.048 |

Not a near-thing: minimum margin is positive in every pair (post-hoc margin analysis,
out/posthoc_margins.json, labeled non-preregistered). The identity rate SATURATES, which
kills P2b's gradient test at that metric — but the predicted script/typology gradient IS
visible in the graded readouts: self-corr deu 0.872 > arb 0.852 > tur 0.842 > jpn 0.791 ≈
zho 0.778, same ordering as per-sentence Jaccard below.

## (b) Parallel-text firing overlap (per-sentence top-32 Jaccard, median; null = max of 20 derangement medians)
| pair | raw | null max | centered | | atom-SET Jaccard (fire≥20, raw) |
|---|---|---|---|---|---|
| en↔deu | **0.333** | 0.016 | 0.362 | | 0.563 |
| en↔arb | **0.306** | 0.016 | 0.306 | | 0.526 |
| en↔tur | **0.280** | 0.016 | 0.306 | | 0.523 |
| en↔jpn | **0.208** | 0.016 | 0.231 | | 0.426 |
| en↔zho | **0.208** | 0.032 | 0.231 | | 0.470 |

All 5 pairs beat the null by 6–21× (P1). Reading (a)+(b) together: the atoms MEAN the same
things across languages, yet meaning-equivalent sentences only co-fire ~a quarter to a third
of their top-32 — consistent with 045's diffuseness result (top-32 is a sliver of ~2450
positive atoms; two near-identical z can cut the ranked mass differently) plus genuinely
language-specific atoms (d).

## (c) Reconstruction: the English-corpus SAE ports across languages, and language ≈ mean offset
| lang | FVU raw | FVU centered | Δ |
|---|---|---|---|
| eng | 0.7144 | ≡0.7144 | — |
| deu | 0.7399 | 0.7124 | 0.0275 |
| tur | 0.7453 | 0.7127 | 0.0326 |
| arb | 0.7476 | 0.7109 | 0.0367 |
| zho | 0.7937 | 0.7227 | 0.0710 |
| jpn | 0.7972 | 0.7165 | 0.0807 |

Raw penalty vs English is small (0.026–0.083, ordered by script/typology distance) and
centering (z_L − μ_L + μ_eng; offsets 0.18–0.30 of ‖z‖) removes essentially ALL of it —
FVU_centered 0.711–0.723 ≈ English 0.714 (P3: ≥0.03 in 4/5, deu just under at 0.0275).
062's "language identity ≈ per-language mean offset" holds in reconstruction space too.

## (d) Language-exclusive atoms (fire ≥20 in exactly 1 of 6 langs)
Raw: union 918 atoms, **37.9% exclusive**. Centered: union shrinks to 557 (−39%),
exclusive 19.4%. So roughly half the apparent language-specificity (and a large share of
union membership itself) is mean-offset-driven firing — but a real ~19% language-exclusive
core survives centering. P4 asked exclusive ≥10% raw (yes) AND at-least-halving
(0.1939 vs threshold 0.18955): **missed by 0.004** — scored FALSE as frozen; direction right.

## Predictions → Brier (frozen probabilities)
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 pair-Jaccard beats null, all 5, raw | 0.90 | **TRUE** (min 6.4×) | 0.0100 |
| P2 identity ≥0.5 all pairs (centered) | 0.75 | **TRUE** (all 1.000) | 0.0625 |
| P2b identity en↔deu > en↔jpn | 0.60 | **FALSE** (1.000 = 1.000, ceiling tie) | 0.3600 |
| P3 centering buys ≥0.03 FVU in ≥4/5 | 0.60 | **TRUE** (4/5; deu 0.0275) | 0.1600 |
| P4 exclusive ≥10% raw AND halves centered | 0.55 | **FALSE** (0.1939 > 0.18955 by 0.004) | 0.3025 |

**Mean Brier = 0.179.** P2b's miss is metric saturation, not a wrong theory — the gradient
it predicted appears in self-corr and Jaccard orderings (deu closest, jpn/zho farthest).
P4 was an honest near-coin-flip that landed a hair on the other side of "halves".

## Limitations / what did NOT run
- Nothing truncated; all preregistered cells + gates ran. Post-hoc additions: margin
  analysis and raw-norm/offset table (labeled; no prediction scored on them).
- Per-language SAE training (arm a) pre-registered OUT OF SCOPE (2009 rows vs 280k
  training samples; 044 identifiability failure) — cross-TRAINING atom universality is
  untested here; this row certifies cross-INPUT-language universality of one dictionary.
- Matcher candidate set = atoms firing ≥20 in BOTH languages (~270-300 of 16384) — identity
  is certified for the shared frequent core only; the 19% exclusive atoms are by
  construction outside the matcher. Rare-atom identity unknown (and per 044, rare atoms
  are the seed-idiosyncratic ones).
- FLORES = news-ish domain, 2009 rows; single SAE (h16384 k32 s0), single encoder (SONAR),
  Pearson profiles, 20 derangements. FVU ~0.71-0.80 regime throughout (the dictionary
  misses most z content per 042-045; "atoms mean the same thing" ≠ "atoms carry the
  sentence"). T3-exploratory.

## Follow-up worth funding? **Y (narrow).**
1. The graded matcher (margin / profile-corr instead of saturated argmax-identity) is now
   the right instrument for the typology gradient AND for re-running 044's gated novelty
   cell (activation-correlation across widths/seeds) — data already on the box.
2. The centering-resistant ~19% exclusive core (557-atom union): are these script/tokenizer
   detectors or content? Cheap: top-activating FLORES sentences per exclusive atom.

## Provenance / hygiene
Box .../campaign100/046-crosslingual-atoms/ (z npz kept on box, 40 MB; no ckpts created;
FLORES tarball deleted after extract). out/: results_046_{smoke,full}.json,
posthoc_margins.json, run.log, DONE. GPUs left idle; tmux c100_046 killed (ours only).
Local commit, no push.
