# RESULT — 049 atom-paraphrase-invariance: do atoms fire invariantly across paraphrases, or are they surface-bound?

**Tier: T3-exploratory.** GPU CVD=0→phys0 only (guarded, 35 MiB pre-claim), tmux c100_049,
smoke passed first, full chain 165 s (prep 10 s + SONAR-encode 12×5000 = 135 s + analyze 20 s),
DONE 06:29:22Z, SELF-HARVESTED in-session 2026-08-03. Prereg frozen before run; data
(PAWS labeled_final + GLUE QQP via HF) verified on box before freeze. n=5000 pairs × 6 cells.

**Verdict: ★ BOTH, quantified — atom co-firing is genuinely semantic AT MATCHED SURFACE, but
wording moves it about as much as meaning does, and per-atom profiles are topic-bound rather
than proposition-bound.** Three headline numbers: (1) paraphrases co-fire a median 0.561 of
their top-32 vs 0.333 for PAWS non-paraphrases at the SAME lexical overlap (0.88 both; para >
nonpara in 10/10 overlap-matched deciles) — meaning matters, decisively. (2) But same meaning
in different words (QQP duplicates, lowest-overlap tercile, lex 0.27) co-fires only 0.208 —
numerically AT THE FLOOR of 046's translation band [0.208, 0.333], while different-meaning/
shared-words pairs sit at its ceiling (0.333). Rewording within a language costs the same
co-fire as translating to Japanese; 046's "diffuse selections" result is about SURFACE
REALIZATION, not about crossing languages. (3) Per-atom, P4 FAILED: activation profiles are
almost as correlated across non-paraphrase pairs (median 0.836) as across paraphrase pairs
(0.906) — semantic index median 0.069 < 0.10 — because PAWS non-paraphrases share topic, and
per 047 atoms are topic detectors. Atoms are meaning-indexed at the topic/lexical-field level,
not at the propositional level. Mean Brier **0.179** (P1 T, P2 T, P3 T, P4 F).

## Gates (all PASS, full n)
- **G1 ckpt identity**: pile-4k FVU **0.7312** (band 0.731±0.02; ≡045/046/047).
- **G2 overlap control**: median lexical overlap paws_para 0.8824 vs paws_nonpara 0.8889 —
  gap **0.007** < 0.10. PAWS design premise holds; absolute-gap clause of P1 valid.
- **G3 null discrimination**: scrambled-pairing (20 derangements) median J32 0.016–0.049 in
  all cells, < observed median in every content cell (nonpara 0.333 vs null-max 0.032 = 10×;
  paws_rand ≈ its null 0.0159, as it should be).
- **G4 embedding + norm sanity**: median z-cos 0.958 (para) > 0.861 (nonpara) > 0.112 (rand);
  mean ‖x_n‖ cell ratio 1.349 < 5.

## (a) Per-cell co-fire (median over 5000 pairs; J32 = top-32 Jaccard, wJ = mass-weighted)
| cell | meaning | surface (lex med) | J32 | wJ | z-cos | null J32 med |
|---|---|---|---|---|---|---|
| paws_para | same | 0.882 | **0.561** | 0.571 | 0.958 | 0.016 |
| paws_nonpara | diff | 0.889 | **0.333** | 0.375 | 0.861 | 0.016 |
| paws_rand | diff | 0.056 | 0.016 | 0.032 | 0.112 | 0.016 |
| qqp_dup_high | same | 0.700 | 0.391 | 0.469 | 0.892 | 0.032 |
| qqp_dup_low | same | 0.267 | **0.208** | 0.262 | 0.687 | 0.049 |
| qqp_rand | diff | 0.043 | 0.032 | 0.076 | 0.182 | 0.032 |

Rough additive decomposition: meaning effect at matched surface = 0.561−0.333 = **+0.228**
(PAWS); surface effect at matched meaning = 0.391−0.208 = **+0.183** (QQP terciles).
Comparable magnitudes — the top-32 selection is a roughly even blend of what is said and how.
Overlap-matched deciles (pooled lex deciles with ≥100 pairs/side): para > nonpara **10/10**,
per-bin gap +0.12 to +0.24, no trend with overlap — the meaning effect is not a residual
overlap artifact. Length control: corr(J32, pair length) −0.014 (para) / +0.133 (nonpara) /
−0.404 (qqp_dup_low — longer reworded questions co-fire less; consistent with more room to
reword).

## (b) Baseline: does the sparse code keep the embedding's separation? (AUROC, para vs nonpara)
z-cos **0.775** > wJ 0.760 ≈ J32 0.760 >> lexical overlap **0.491** (≈chance — PAWS's
adversarial construction confirmed on our instrument). The raw embedding separates PAWS only
modestly (it is an adversarial set), and the k=32 code retains ~98% of that AUROC. The
dictionary is not adding semantic separation over z, but it is not destroying it either.

## (c) Per-atom: invariance vs lexical-boundness (n_cand=414 atoms, fire≥50 both sides; BoW 155 mid-freq words)
- inv_para (Pearson across paraphrase pairs) median **0.906** [q25 0.877, q75 0.959] — higher
  than every 046 cross-language self-corr (deu 0.872 best). Frequent atoms are highly stable
  under rewording *in the profile sense*.
- BUT inv_nonpara median **0.836**: profiles correlate almost as well across pairs that mean
  DIFFERENT things but share topic/words. Semantic index (inv_para − inv_nonpara) median
  **0.069** (P4 predicted ≥0.10 → FALSE); 41.8% of atoms ≥0.1, 41.3% "surface-like" (<0.05).
- Lexical-boundness is weak at the single-word level: best-word |corr| median 0.116; **0/414
  atoms** have best-word corr > inv_para; corr(inv_para, bow_best) = 0.25. A few clear
  word/topic atoms exist (atom 15428 "film" bow 0.727, inv_para 0.963, sem_idx ≈0 — a pure
  topic detector). Most-semantic atoms (sem_idx +0.18–0.20, e.g. 10645, 9108, 3246) keep high
  inv_para ~0.85 while dropping to ~0.66 on non-paraphrases — a minority (~40%) of atoms carry
  meaning-beyond-topic.
- Read with 047: atoms are topic/lexical-field detectors; topic survives both paraphrase AND
  PAWS-style meaning flips, so per-atom "invariance" is mostly topic persistence.

## (d) Tie to 046 (translation band [0.208, 0.333], FLORES)
Same-language paraphrase with high overlap: 0.561 — well ABOVE the band (P3). Same-language
paraphrase with low overlap: 0.208 — AT the band floor (jpn/zho level). High-overlap
different-meaning: 0.333 — at the band ceiling (deu level). The translation co-fire deficit
is reproduced within English by rewording alone; conversely shared wording buys as much
co-fire as full translation-equivalence of meaning. (Cross-domain caveat: QQP questions vs
FLORES news; lengths differ ~9 vs ~30 words.)

## Predictions → Brier (frozen probabilities)
| pred | P | outcome | Brier |
|---|---|---|---|
| P1 para−nonpara gap ≥0.05 AND ≥8/10 matched deciles | 0.80 | **TRUE** (0.228; 10/10) | 0.0400 |
| P2 nonpara ≥5× its scrambled null | 0.85 | **TRUE** (21×) | 0.0225 |
| P3 paws_para median J32 > 0.33 | 0.70 | **TRUE** (0.561) | 0.0900 |
| P4 median per-atom (inv_para − inv_nonpara) ≥ 0.10 | 0.75 | **FALSE** (0.069) | 0.5625 |

**Mean Brier = 0.179.** P4's miss is informative, not instrumental: the profile-correlation
instrument saturates on topic overlap (PAWS nonpara pairs are topic-matched by construction),
so proposition-level per-atom selectivity is smaller than predicted — consistent with 047's
topic-detector reading and worth stating as the finding.

## Limitations / what did NOT run
- All preregistered cells, gates, and analyses ran; nothing truncated. No post-hoc analyses added.
- PAWS non-paraphrases share topic and most content words — the "meaning effect" (+0.228) is
  meaning-beyond-topic-and-wording (structure/roles/relations), not meaning-vs-unrelated.
  Conversely the per-atom semantic index is conservative for the same reason.
- The para+low-overlap cell comes from QQP (different domain, question style, shorter); the
  band comparison to FLORES translations crosses domain and length regimes. In-domain
  low-overlap paraphrases (e.g. backtranslated wiki) would be the clean 4th cell.
- QQP "duplicate" labels are noisy (crowd-sourced); tercile split inherits that noise.
- Single SAE (h16384 k32 s0), single encoder (SONAR), FVU ~0.73 regime — statements are about
  the dictionary's top-32 selections, which carry a minority of z (042-045). J32 uses top-32
  sets; FIRE_MIN 50/5000 rate-matched to 046. T3-exploratory.

## Follow-up worth funding? **Y (narrow).**
1. The 40%-semantic / 40%-surface per-atom split: rerun semantic index with PAWS pairs
   stratified by shared-content-word count, or with role-swap minimal pairs (061/062 battery)
   to isolate proposition-level atoms from topic atoms — directly tests whether ANY atom
   encodes structure rather than field.
2. In-domain low-overlap paraphrase cell (backtranslation of the same PAWS sentences) to
   confirm the "rewording ≡ translation" equivalence without the QQP domain confound.

## Provenance / hygiene
Box .../campaign100/049-atom-paraphrase-invariance/ (231M: z npz + cells json + per-atom npz
kept on box; no ckpts created; PAWS gcs tarball 404'd and removed, HF parquet used). Repo out/:
results_049_{full,smoke}.json, run.log. GPUs left idle (35/15/15 MiB); tmux c100_049 killed
(ours only; server empty after). Local commit, no push.
