ai gen
This is the Campaign100 archive (6 September snapshot). The new research plan has its own current kanban, expanded protocols and multilingual results. Ledger confidence values below are historical heuristic aggregates, not calibrated probabilities of scientific truth.

Campaign100 · what is done, what is not

One hundred pre-registered experiments on what a SONAR sentence embedding z encodes (2026-08-01 → 08-08), plus the consolidation pass that hardened the top three results (08-14). Every row ran. What remains is promotion, write-up, and the follow-up backlog. Board built 2026-09-06 from board_data.json.
100
recorded experiment rows
9
tracked claims
10
program tasks

1 · Program-level board the campaign as a whole

Three columns, ten items. The two done items are the campaign itself and the consolidation pass. Everything else is what stands between the results and a paper.

Done

2

100/100 rows run

Every row has PREREG_LITE.md + RESULT.md + a Brier score; campaign closed 2026-08-08.

Consolidation H1–H3

All three hardening threads DONE 2026-08-14 with codex review; one real bug found and fixed in each of H2/H3, H1 crux confirmed.

In progress

1

Paper draft

PAPER_DRAFT.md started 2026-09-06 from CAMPAIGN_FINAL + ledger + consolidation. Text drafted; figures not made; related work thin.

Not done

7

Ledger rerun

Rerun bayesian_ledger.py with H1/H2/H3 as support rows and 061/062 demoted contra→refuted (CONSOLIDATION.md names this as the clean next step).

Formal promotion

No campaign claim has been formally promoted above T3 in CLAIMS_LEDGER.md. C1 is promotable (H1 breaker passed); the ledger entry has not been written.

Figures

No publication figures exist for the campaign (only the z-explorer artifact and status.html). Needs ~6: binding null + power curve, capacity knee vs R(D), operator algebra, fabrication ROC, ontogeny, ledger.

Human calibration (060)

Designed and proxy-piloted with an LLM panel; never run with humans.

TAE-Bench public release

Packaged and validated locally (row 100); not pushed to any public remote (protocol: no git push). Needs a repo, license, and a GPU smoke test on a fresh clone.

Blocked rows

009 (needs a binding teacher) and 051 (needs per-embedder battery recalibration) remain unresolved.

Follow-up backlog

63 rows flagged a follow-up worth funding; only the three consolidation threads have been executed. The rest are unfunded.

2 · Claims board C1–C9 from the Bayesian ledger

Posterior = the ledger's own credence (T3 bookkeeping, not a promotion). A claim moves right only when an independent breaker passes. Note the empty formally promoted column: nothing has been written into CLAIMS_LEDGER above T3 yet, even where the breaker already passed.

Exploratory (T3)

5
C2

~460-bit sentence-specific capacity, ~4-proposition knee

none yet

Needs: Freeze the 460 ceiling in a fresh prereg (066 was post-hoc); separate SONAR from ladder evidence; independent breaker.
C9

z leaks propositional / sensitive content (RAG, canaries)

none yet (H2 corrected the monitor choice: nn-cos does not transfer to natural text)

Needs: Second embedder, second seed, non-templated corpus.
C7

‖z‖ encodes length/specificity, not thematic semantics

none

Needs: Decoder scale-invariance across a norm range; cosine-controlled density analysis frozen in a prereg.
C6

SAE atoms semantic at the frequent core, seed-idiosyncratic tail

none

Needs: Cross-seed stability at a matched frequency band; the 044 contradiction must be reconciled.
C4

Order/role code forms mid-training (substrate-specific ontogeny)

none

Needs: Ladder-only evidence; needs a SONAR-side checkpoint series, or reframing as three separate ontogenies (order code / atoms / operators).

Open / half-refuted

2
C8

No universal crosslingual role code

H1 refuted the Japanese exception; the claim moves toward the null but the ledger has NOT been rerun.

Needs: Mechanical ledger rerun with 061/062 demoted; one non-template Japanese/German stimulus set.
C5

Fabrication is off-manifold and gate-catchable

087: near-manifold interpolation fabrication is unrejectable (AUC 0.508). Claim only half-true.

Needs: Rewrite as two claims (gross vs near-manifold); the near-manifold half is open.

Breaker-passed

2
C1

No linear role binding in pooled z

H1 (2026-08-14): English null survives on real SONAR with power certified (linear 0.509, MLP 0.495; planted d=1.0 → 0.973). Both contra rows (061/062) refuted.

Needs: Formal promotion entry in CLAIMS_LEDGER (prereg + breaker already exist); ledger rerun with H1 as a support row.
C3

Closed-class operators are linear, composable, causally steerable

H2: codex certified the steering core; garble-counting bug fixed (negation honest success 0.70 at α1, not 0.90); natural-text over-steer caught by rt_cos (AUC 0.845). latent_rewrite.py built.

Needs: Scope statement in the paper (tense safe; negation narrow α window; vertical frame-bound). Multi-seed run on natural text.

Formally promoted (>T3)

0

3 · Row-level board 100 experiments by degree of done-ness

Every row has a frozen prereg, a RESULT.md and a Brier score, so "ran" is not the interesting axis. Lanes instead track how far each result has travelled: blocked → partial → exploratory (T3) → hardened by a breaker → shipped as a tool. Two rows moved backwards after the campaign (refuted). Click a card for the finding, method, and detail.

Blocked / instrument failure

2
009 · A

Binding distillation

Blocked by its own gate — a 3-billion-parameter LM's pooled sentence representations also fail to bind roles, so there was nothing to distil from.

blocked · Brier 0.022
051 · F

Embedder binding sweep

All six small sentence embedders FAIL the battery's positive-control gate (within-construction ceilings 0.54–0.74 vs SONAR's ~1.0) — so the sweep returns no clean nulls, and one real twist: role information is faint in all of them.

blocked · Brier 0.220

Partial

1
060 · F

The judgment anchor

SONAR's within-topic similarity is not merely role-blind — it is ANTI-meaning: given a paraphrase and a role-swap of the same sentence, the panel picks the paraphrase 90/90 times while z-cosine picks the role-swap 30/30, and within-topic correlation with judged similarity is NEGATIVE (−0.218).

twist · Brier 0.233

Refuted after campaign

2
061 · G

Case-marking languages

The first order-invariant role binding of the entire program appeared here — German 0.668 and Japanese 0.658 on the strict primary cell. Row 062's fresh-vocabulary replication then rewrote the ranking: Japanese held (0.696), German did not (0.551).

signal · Brier 0.154
062 · G

Cross-lingual role transfer

SONAR's one vocabulary-general role code (Japanese, replicated at 0.696 on a fresh lexicon) does NOT transfer to other languages — and German's 061 positive fails fresh-vocabulary replication (0.551). Binding is language-local in z.

twist · Brier 0.163

Exploratory result (T3)

65
007 · A

Two-slot bottleneck

Given two independently attention-pooled slots, the model made them redundant twins rather than factoring roles — each slot carries the full lookup.

twist · Brier 0.216
008 · A

Anti-binding mechanism

No meaningful early anti-binding dip at 10% budget — the below-chance readings are noise-scale, because the surface code has barely formed. The lasting yield was an arithmetic insight.

null · Brier 0.192
010 · A

Case-marked scrambled MT

A scrambled-order case-marked target language was unlearnable at matched budget — 0.503 case accuracy vs 0.99 for fixed slots. The output scaffold was doing the role work all along.

twist · Brier 0.238
011 · B

Pooling-input role code (v4)

A scale-consistent probe of the exact states mean-pooling consumes finds no construction-invariant role code — certified by planted-signal recovery. SONAR's binding null is stack-deep.

null · Brier 0.267
012 · B

Role-consistent attention heads

42 of 384 attention heads route agent→predicate consistently across all constructions — Bonferroni-significant, held-out-replicated, killed by label permutation. Role lives in the computation.

signal · Brier 0.380
013 · B

Role-head knockout

Ablating all 42 role heads leaves reconstruction role-fidelity untouched (+0.001) — the role-structured attention is causally inert. Only surface heads matter.

null · Brier 0.172
015 · B

Pooler retrofit

A trained, transfer-disciplined attention pooler on frozen token states recovers no transferable role code at any depth — certified. C1 is a property of the states, not the readout.

null · Brier 0.092
019 · B

Decoder-layer mirror

Role never crystallizes into a linear code anywhere in the decoder; order is 0.999 early and DECAYS with depth; cross-attention reads z as a length-1 sequence — no copy pathway.

twist · Brier 0.041
020 · B

z-read lens

z enters the decoder as a constant per-layer bias — the cross-attention contribution is bit-identical at every generation step. There is no stepwise "reading" of z.

twist · Brier 0.068
021 · C

Fabrication taxonomy

Semantic fabrication on clean sentences is 3.6% — an order of magnitude below the program's 20–30% headline — length-scaling, entity-dominated, and 65% invisible to a cosine gate.

signal · Brier 0.179
022 · C

Entropy signature

Fabricated spans are high-entropy guesses (Cohen's d 1.20); an entropy detector reaches AUC 0.87 — but it is redundant with cosine, not a complementary signal.

signal · Brier 0.267
023 · C

Off-manifold dose-response

No fail-open band exists off-manifold — the scary "fluent lies" picture collapses. The only potent direction is interpolation toward another real embedding, where fabrication and garble co-rise.

null · Brier 0.388
024 · C

LM-prior mechanism

Fabrication is the decoder's language-model prior showing through where z fails to rescue a prior-implausible token — a mechanism, and a detector feature orthogonal to entropy and cosine.

signal · Brier 0.220
025 · C

Fidelity-gate reliability

The only deployable cosine gate — self-consistency, holding just the latent — accepts a fluent, completely different reconstruction ~90% of the time at every threshold.

signal · Brier 0.116
026 · C

Decoding strategy

Greedy is already at the faithfulness ceiling; beam search does not reduce fabrication; temperature is the dangerous knob, with a cliff between 0.7 and 1.0.

twist · Brier 0.216
027 · C

Round-trip attractors

Iterating encode→decode is a per-sentence near-identity map — 2000 seeds give 2000 distinct fixed points, no consolidation toward generic attractors.

null · Brier 0.168
028 · C

SAE decomposition

A single length ≈ inverse-density difficulty axis governs the sparse-autoencoder decomposition (recon cosine vs length r=−0.78), unifying length, intrinsic dimension, and fabrication.

signal · Brier 0.188
029 · C

Decode-verify auditor (v2)

An NLI decode-then-verify auditor ranks fabrications BELOW chance (AUC 0.478) — it shares cosine's blind spot. And the 3.6%-vs-20–30% gap is input cleanliness, not method.

null · Brier 0.467
030 · C

Confidence calibration

The decoder's token confidence rank-orders fidelity (AUC 0.738) but is systematically UNDER-confident; PAV recovers a proper map; abstention is a coarse lever.

twist · Brier 0.181
031 · D

Local intrinsic-dimension field

The manifold is decisively inhomogeneous — local intrinsic dimension varies 2–3× by length and domain (short 31 < long 69; fiction 36 ≪ news 60). Ordering claimed, magnitudes withheld.

signal · Brier 0.202
032 · D

Geodesic vs linear interpolation

Manifold-following geodesics stay coherent where straight chords collapse — 62.5% coherent vs 50%, only 4% word-salad vs 17%. But geodesics buy coherence, not smooth meaning-blending.

signal · Brier 0.262
036 · D

Norm semantics

The magnitude ‖z‖ encodes specificity/information content (partial r 0.41 with perplexity) — but is not a usable control: scaling z ±30% changes the decode not at all.

twist · Brier 0.249
038 · D

Antipodal decoding

The antipode −z is unstructured — statistically indistinguishable from a random same-norm vector, not an opposite, not a valid latent, orthogonal to the negation operator.

null · Brier 0.151
039 · D

Persistent homology

No robust global topology — no persistent loops or voids beyond a Gaussian blob, and cyclic linguistic attributes do not trace geometric cycles. Best calibration of the campaign.

null · Brier 0.056
040 · D

Whitening robustness

Every audited campaign finding survives a change of basis — and whitening does NOT un-blind the fabrication gate (ZCA: −0.011 AUC), killing the last basis-artifact escape hatch.

signal · Brier 0.312
041 · E

Crosscoders across depth

Dictionary atoms are depth-LOCAL, not persistent: of 13,706 firing atoms in a crosscoder spanning L12, L24, and z, only 26 — 0.19% — are shared across all three.

null · Brier 0.266
042 · E

The residual is structured

The half of z the SAE fails to reconstruct is NOT noise: adding the recovered residual to the decode lifts chrF by +25.3 (a Gaussian control LOWERS it), and its ICA components sit ~69× above the matched-noise null.

signal · Brier 0.264
043 · E

What the dictionary drops

The residual out-informs the reconstruction on every certified probe: token length (R² .907 vs .723), word content (AUC .937 vs .878), domain (.832 vs .782) — the k=32 dictionary keeps a MINORITY of the linearly-decodable content.

signal · Brier 0.106
044 · E

Width buys nothing

Across a 32× width sweep at matched budget, feature splitting is essentially absent (4.1% → 0.4%, shrinking with width), FVU is flat from h=2048 to h=65536, and the residual beats the reconstruction on all five probe families at every width.

null · Brier 0.144
045 · E

The k-atom curve

Keeping more atoms makes L2 reconstruction WORSE past m=64 (FVU U-curve, 0.69→0.83) while decode quality keeps rising monotonically (chrF 21→35) — the dictionary's error metric and the decoder disagree about what matters.

twist · Brier 0.146
046 · E

Atoms are meaning-indexed

The English-trained dictionary's atoms fire on MEANINGS, not English: activation-correlation identity is 1.000 across all five language pairs (German, Japanese, Turkish, Chinese, Arabic), raw and centered.

signal · Brier 0.179
047 · E

Frames vs atoms

Semantic frames explain only ~3.3% of z's variance and the frame-supervised basis loses to plain PCA-50 — yet 46 of 50 frames have an atom firing selectively for them. Atoms are topic and lexical-field detectors, not frame-structure detectors.

twist · Brier 0.201
048 · E

Atom ontogeny

Atoms crystallize gradually and front-loaded: half the final inventory is matchable at 5.7% of the training budget, 98% by 74%, with no late reorganization — and ~31% of top-frequency atoms exist already in the UNTRAINED encoder.

signal · Brier 0.179
049 · E

Paraphrase invariance

Atom co-firing is genuinely semantic — paraphrases beat equally-word-matched non-paraphrases in all ten surface-overlap deciles — but rewording alone moves co-firing as much as meaning does, reproducing 046's translation deficit inside English.

twist · Brier 0.179
050 · E

Dead-feature necropsy

The row's premise is empty, and that is the result: no atom ever dies in training, the aux-k revival mechanism never engages (aux loss ≡ 0 across all 8 runs, 320 epochs), and 'dead' atoms are ordinary rare-topic atoms that don't transfer across an eval-distribution shift.

null · Brier 0.145
052 · F

LASER vs LaBSE: the objective

It's the generative MT decoder: 45M-parameter LASER lands in SONAR's probe-power regime while 471M LaBSE patterns with the small contrastive models — training to GENERATE translations installs the surface code that training to RANK them does not.

signal · Brier 0.219
053 · F

Instructions are inert

Telling an instruction-tuned encoder 'represent this sentence by who performs the action' does nothing: the ceiling shifts by +0.004–0.006 — indistinguishable from a scrambled-prefix control — and opposite instructions produce embeddings at cosine 0.999.

null · Brier 0.092
054 · F

The concept-space LM

Even an LM that PLANS in SONAR space develops no binding: all three depths of a 1.6B Large Concept Model stay in SONAR's exact probe regime (max ceiling shift +0.016), the binding cell is dead chance throughout, and the surface-order code is mildly amplified, not replaced.

null · Brier 0.132
055 · F

Continuous thought

Reasoning pressure doesn't install binding either: across all six of Coconut's continuous-thought steps the binding cell stays at chance (0.494–0.515), and the within-ceiling DECAYS as thoughts progress. Five pressures tested across the arc, five NOs.

null · Brier 0.243
056 · F

Cross-attention doesn't bind either

The folklore that rerankers 'handle roles via cross-attention' is wrong: on surface-incongruent role swaps the cross-encoder scores 0.468 — BELOW chance, actively preferring the wrong candidate whose word order matches the query — and there is no role layer to localize.

twist · Brier 0.281
057 · F

Diffusion latents

Iterative denoising is the seventh NO: role binding sits at chance at every noise level (32 of 32 cells), nothing ever 'commits' to who-did-what across the denoising trajectory, and what does commit early is the surface scaffold — function words recover 4× faster than content words.

null · Brier 0.134
058 · F

Speech is text with an accent

SONAR's speech encoder lands spoken sentences essentially on top of their transcripts (cosine 0.912, retrieval P@1 0.967 even among role-swap twins), the text decoder reads speech-z near-natively with zero fabrication — and the binding null extends to a second modality, the arc's eighth NO.

signal · Brier 0.166
063 · G

The capacity tax

Languages pay a LEVEL tax, not a knee tax: round-trip fidelity ranges from 89.7 (English) down to ~35 (Chinese/Japanese) at every length, but the capacity knee sits at ~2 sentences of content everywhere — and the language offset is not load-bearing for the decoder.

twist · Brier 0.154
064 · G

The language vector

Translation is NOT a constant offset and the language vector is causally near-inert: v_lang explains only 9–13% of per-sentence translation displacement, steering flips the output language 0% of the time, and a plain decoder-token switch already delivers 80–88% of the translation ceiling.

null · Brier 0.210
065 · G

Code-switching monolingualized

The matrix language owns a mixed sentence (one embedded word moves z by ~2% of the between-language displacement), interpolating between translations has NO code-switched region at any point, and the round-trip actively monolingualizes — insertions get translated into the decoder token's language.

twist · Brier 0.220
066 · H

The bits budget

One SONAR vector carries ~180–195 bits per sentence in EVERY language — Japanese's terrible round-trip fidelity nearly vanishes in bits (163 vs English 182): the information is in z, the free-running decoder just can't surface it. Specific information caps at ~0.45 bits per dimension.

twist · Brier 0.295
067 · H

The dimension ladder

Capacity scales with bottleneck dimension up to the architecture's rank-256 ceiling and saturates exactly there — specific bits climb 110 → 319 → 404 across d_z 16→64→256, while the d_z=1024 reparameterization placebo is dead flat (knee excess 0.35, within the seed gap of 0.37).

twist · Brier 0.222
068 · H

The budget is in bits

The capacity knee is denominated in BITS, not tokens: across natural tiers the bits-at-knee is constant to 1.08× (352/327/347) while random text — the one genuinely high-perplexity tier — knees at HALF the token length of predictable text, exactly the leftward shift a bits-budget predicts.

signal · Brier 0.224
069 · H

Conjunction is subadditive

Storing 'A and B' always costs less than A plus B — even unrelated clauses compress ~8%, paraphrase pairs compress 28%, and when demand exceeds the ~460-bit ceiling the shortfall splits evenly across both clauses: graceful degradation, not clause dropout.

signal · Brier 0.130
070 · H

Numbers don't cliff

Random numbers survive to 12 digits intact in a short sentence (exact-match ≥0.96) — a 12-digit number is only ~40 bits, trivially under the ~460-bit budget. Digits aren't fragile; the cliff is purely capacity pressure, appearing only near the knee (exact drops to 0.31).

twist · Brier 0.195
071 · H

The entity ceiling

About three distinct rare entities survive a round-trip at fixed length — but it's capacity, not slots: give the sentence more length and it sustains five or six. Each rare entity costs 4–6× a number, and the first is bit-protected while later ones pay.

twist · Brier 0.330
072 · H

What gets deleted first

Pushed past the knee, SONAR deletes the DIRECT OBJECT first and keeps time and quantity last — the reverse of gist-over-detail. The predicate skeleton is NOT preserved as a unit: its object collapses into a repeated placeholder while the subject survives.

twist · Brier 0.395
073 · I

When the order code forms

SONAR's surface-order code and its anti-transfer signature form gradually in mid-training (onset ~11% of budget, saturated by ~74%) — not an early cheap heuristic, not a late abrupt phase transition, and later than the dictionary atoms crystallize.

null · Brier 0.276
074 · I

Grokking: none

15,000 epochs of over-training — 5× the standard budget — induce ZERO delayed emergence: role abstraction stays at chance at every one of 15 checkpoints, the order code freezes early and weak, and validation peaks then declines. Compute alone buys no abstraction.

null · Brier 0.120
075 · I

Anatomy up to rotation

Across 8 seeds, SONAR's anatomy is stable up to a global ROTATION of the z-basis: operator directions are seed-random (raw cross-seed cosine ≈ 0) but snap to 0.83 after a single orthogonal alignment, while the knee, order code, and reconstruction are near-identical (CV ≤ 0.03).

signal · Brier 0.104
082 · K

The cosine gate's blind spot

A cosine-similarity monitor on z cannot both admit paraphrases and reject meaning flips: a fluent negation costs less cosine (0.069) than a meaning-preserving synonym swap, so no threshold separates them — and closing the hole entirely requires τ≈0.98, which rejects nearly everything.

twist · Brier 0.334
083 · K

Covert capacity, and its catch

A z-vector can hide ~103 bits per sentence that slip past BOTH a cosine monitor and a decode-parse monitor undetected (recovered losslessly) — an order of magnitude more than expected — but those payloads are NOT invisible: a manifold-distance monitor catches the directional ones at AUC 0.92–0.94 and a norm check catches the radial one at 0.998.

twist · Brier 0.207
084 · K

Canaries survive retrieval

Role structure survives a SONAR-RAG round-trip cleanly: a role-correct query retrieves the right proposition over its role-swapped twin 48/48 times, a pure in-store role-swap NEVER wins top-1 (false-match rate 0.000), and retrieved embeddings decode back to the exact canary.

signal · Brier 0.207
085 · K

One word breaks RAG

A single shared surface word overturns SONAR's role-binding margin: it buys +0.123 cosine to a role-WRONG distractor, so ~2 shared words fully overpower the +0.159 role-correct margin — and when the distractor wins, downstream QA is wrong 100% of the time, returning the asker's own name as the answer.

signal · Brier 0.099
086 · K

The ensemble monitor

An ensemble of six detectors catches all seven corruption types at AUC 1.000 (0% false positives) when it can insist on a tight re-encode — but that also rejects 82% of benign paraphrases; relax it to admit paraphrases and the blind spot returns, now a covert null-carrier payload that no geometric detector sees.

twist · Brier 0.182
087 · K

Fail-closed decoding

A density gate is a near-perfect fail-closed decoder against gross off-manifold fabrication (reject-ROC 1.000, keeps clean AND the faithful covert channel) — but it is blind to interpolation fabrication: a midpoint of two embeddings sits at clean-level density yet fabricates 92.5% of the time.

twist · Brier 0.330
088 · K

Membership is memorization

An embedding does not inherently betray whether its sentence was in training: a normally-trained model leaks membership at AUC 0.546 (chance), an untrained model at 0.522 — the signal only appears (AUC 1.000) when the model has memorized, and it rides reconstruction fidelity.

null · Brier 0.088
089 · K

Style-transparent, identity-opaque

z is a style encoder: register (0.998), topic (0.963), sentiment (0.903), and formality (0.873) are all highly linearly readable — but demographic identity is not, collapsing to chance once topic is partialled out, and z leaks identity LESS than the raw text would.

signal · Brier 0.186
093 · L

Adversarial stimuli evolution

A codex loop evolving stimuli to break the binding probe found ZERO genuine breaks in 24 candidates: every well-formed, plausible construction leaves the verdict flat, and the only 'breaks' are artifacts the canonicalization checker correctly rejects.

null · Brier 0.238
098 · L

The knee from first principles

Rate-distortion theory predicts the capacity knee with zero free parameters: one source rate (4.5 bits per token) and one channel capacity (460 bits) place all three measured knees — 2 sentences, ~70 tokens, and the mysterious '15.79' — on a single R(D) curve, each at its own distortion bar.

signal · Brier 0.084

Hardened by breaker

23
001 · A

Role-swap contrastive

An InfoNCE objective separated role-swapped pairs on its training vocabulary yet gained nothing on held-out transfer — the objective was satisfied by lexical shortcuts.

null · Brier 0.218
002 · A

QA-head dose-response

A supervised who-is-the-agent head reached 100% held-out accuracy at every dose, while the transfer battery stayed flat at chance.

null · Brier 0.166
003 · A

Structured decoder

Forced to emit an explicit role tuple, the decoder retrieved the agent at 0.998 on order-swapped sentences — yet transfer was exactly chance. The role code is a lexical lookup.

null · Brier 0.062
004 · A

Shuffled-input DAE

Word-shuffling forced order information into z — the model could reorder a scrambled bag (Kendall-τ 0.45) — yet the linear surface-order probe read order worse than baseline.

twist · Brier 0.180
005 · A

Passive curriculum

Making surface order perfectly uninformative about role in 40% of the data — with no role objective at all — induced no transferable role code.

null · Brier 0.078
006 · A

Lexical-diversity phase transition

Holding the objective fixed and only widening the filler vocabulary, role abstraction EMERGES between 117 and 1000 fillers — novel-word retrieval jumps 0.00 → 0.05 → 0.73.

signal · Brier 0.306
014 · B

Swap-patching localization

Patching filler-token states from the role-swapped sentence flips the decode at 0.95 through layer 22, then 0.00 at layer 24 — a flat-then-cliff commitment, identical in actives and passives.

signal · Brier 0.193
016 · B

Nonlinear / kernel probes

Kernel, focal-conditioned MLP, and bilinear probes all null at certified sensitivity — five probe families now agree the states hold no transferable role code, only a perfect surface code.

null · Brier 0.256
017 · B

Position × layer heatmap

The order code is strictly filler-local at every depth (0.86–0.96), never spreading to summary positions — refuting the broadcast-before-pooling hypothesis.

twist · Brier 0.129
018 · B

Function-word carriers

Grammar tokens ("by", "'s", "that") carry no vocabulary-transferring role code — conduits, not stores. Best-calibrated experiment of the campaign.

null · Brier 0.035
033 · D

Analogy operator battery

Tense, number, negation, and question are clean linear, invertible, composable offsets — but voice produces passive form with 0/40 argument swaps. An independent confirmation of the binding gap.

signal · Brier 0.254
034 · D

Causal negation operator

A single linear direction is a causal, dose-responsive, content-specific negation operator that generalizes templated→natural. Best-calibrated ★ of the campaign.

signal · Brier 0.138
035 · D

Operator curvature

The offset operators are a curved field — regional directions rotate 20–25° vs a 3° noise band — but the curvature does not break application: a global offset is good enough for coarse transforms.

twist · Brier 0.179
037 · D

Anisotropy audit

A robustness certificate — SONAR z is nearly ISOTROPIC (mean pairwise cosine 0.068, top PC 3.3%), unlike BERT/GPT, and all four headline claims survive whitening.

signal · Brier 0.193
059 · F

Scale moves nothing

A 44× scale sweep within one fixed architecture and objective — GTR-T5 from 110M to 4.8B, all four scales — leaves every preregistered cell flat: probe ceiling 0.601→0.591, binding at chance throughout, and the possessive code scale can't buy. The ninth NO closes the capacity axis.

null · Brier 0.086
076 · I

Curriculum splits the anatomy

Training ORDER governs reconstruction but not structure: at fixed seed, monotone curricula wreck fidelity (val F1 0.11–0.52 vs 0.645 random, order-variance 104× the seed band) via end-of-training forgetting — yet the surface and relational codes are order-robust (within 2× seed).

twist · Brier 0.340
077 · I

Objective hysteresis

Switching the training objective mid-stream reveals a split: reconstruction follows the LAST objective (recency), but the surface-order code is imprinted by the FIRST objective and never overwritten — a formed code sticks even when the objective that formed it is replaced.

twist · Brier 0.421
078 · J

The operator zoo

The real boundary isn't grammar-vs-semantics: closed-class MARKER swaps all linearize perfectly (before/after, above/below, in/out, near/far, bigger/smaller — success 1.00), while ARGUMENT swaps hit the exact voice binding wall (0/74). A linear offset can substitute a word; it cannot reverse who-relates-to-whom.

signal · Brier 0.125
079 · J

Operators compose

The linearizing operators form a real vector algebra: summing two single-operator directions equals the directly-fitted double transform (mean additivity 0.982, compose-minus-direct gap exactly 0.000), it works across families, and it commutes perfectly.

signal · Brier 0.139
080 · J

Steering is causal

The operator directions don't just describe attribute geometry — they install it: adding α·v steers negation, tense, number, and a spatial marker from 0 to ~1.0 success by α=1, with ZERO off-target collateral, perfect invertibility, and a norm-matched random push completely inert.

signal · Brier 0.158
081 · J

When operators install

Operators have two separable substrates: the offset DIRECTION is present at initialization (a token-embedding fact, negation cos 0.93 and vertical 0.999 at step 0), while CAUSAL usability is learned, decoder-gated, and installs early and abruptly (step 2000–4000) — in a strict order, with the tense operator never installing at all.

twist · Brier 0.255
096 · L

Why bag codes win

A toy mean-pooled autoencoder reproduces the whole binding story — bag-lookup beats abstract role below a critical vocabulary, abstraction emerges above it — but the numerical model FALSIFIES the predicted cause: the transition isn't set by data or capacity, it's set by the surviving order-channel's signal-to-noise.

twist · Brier 0.136
097 · L

What survives pooling

Formally, mean-pooling contextual states smuggles order through ONE channel: the bag reweighted by each token's average received attention — and word order survives only via the relative-position bias, whose antisymmetric part is the sole who-came-before-whom discriminator, exactly 096's β residual.

signal · Brier 0.121

Shipped artifact

7
090 · L

The probe-power kit

A reusable, validated tool that packages the campaign's most important instrument lesson — certify probe power with a planted signal before trusting a null — with a PASS/FAIL certificate that catches the exact uncertified-probe trap that produced the 051 and 062 false nulls.

signal · Brier 0.115
091 · L

The canonicalization checker

A property-based checker that certifies a labeled stimulus set's slot and label conventions before a probe trusts them — and it catches exactly the real bugs rows 017 and 020 hit, each injected bug flipping only its target property with no false alarms.

signal · Brier 0.005
092 · L

The norm linter

A pre-flight linter for any difference-of-means or offset construction that flags magnitude confounds — and building it surfaced a genuine refinement: norm confounds come in TWO modes, artifact (dies under cosine) and confound (norm-defined but survives cosine), needing different fixes.

signal · Brier 0.031
094 · L

The prereg engine

The discipline the whole campaign ran on — frozen predictions, gates, verdict, Brier — extracted into a tested library that reproduces past rows' Brier scores exactly and makes cheating structurally impossible: a certificate hash rejects any post-hoc edit to the predictions.

signal · Brier 0.003
095 · L

The Bayesian ledger

An auditable posterior tracker that turns 95 scattered results into aggregate credences: no-linear-role-binding rises to 0.99, operators-compose to 0.96, the 460-bit capacity to 0.89 — while the genuinely mixed claims (multilingual binding, off-manifold fabrication) correctly stay near 0.5.

signal · Brier 0.042
099 · L

The z-explorer

A self-contained interactive artifact that lets anyone step through the campaign's five headline phenomena — operator steering, interpolation winner-take-all, capacity overload, fabrication with live monitor flags, and the binding demo — every decode a byte-for-byte recorded campaign output.

signal · Brier 0.035
100 · L

TAE-Bench release

The campaign's reusable assets ship as one runnable, self-certifying benchmark: binding battery + five instrument kits + honesty gates + checkpoint manifest + ledger, with a single entry point that reproduces the flagship no-binding null on CPU, offline.

signal · Brier 0.132

4 · Follow-up backlog rows that flagged "follow-up worth funding: Y"

Sixty-three rows asked for a follow-up. Only the ones marked covered have had one executed (via consolidation H1/H2/H3). The rest are unfunded ideas, listed with the row's own words.

001

Role-swap contrastive

Follow-up worth funding? Yes, conditionally — 002's dose-response (running) is the direct next probe; if λ=1 aux-QA also fails to transfer, the "shortcut satisfaction" story strengthens and 096 (toy theory) should formalize it.

009

Binding distillation

Follow-up worth funding? Y (redirected). Not "distill from a pooled teacher" (dead — nothing binds under the pool). The live questions this surfaces: (1) does ANY LM state bind roles *before pooling* — probe last-token / per-content-token reps of Qwen2.5-3B on the battery (a mini-051-for-decoders); if yes, distill from the UNPOOLED signal (cross-attention or per-token) rather than a single pooled vector. (2) 059 GTR-T5 base→XXL scale sweep is now the clean way to ask if scale ever induces *pooled* binding — 009 con

019

Decoder-layer mirror

Limitations / follow-up - Teacher-forced confound (headline, pre-registered): order is trivially in the input, so this localizes where order/role are linearly+transferably decodable in the decoder residual stream, NOT order *recovery from z alone*. The genuine recovery question needs FREE-RUNNING generation with generated-token alignment — the clean follow-up (worth funding: Y — it would test whether the decay-with-depth reverses when order must be reconstructed rather than read). - Self-attn attribution mean-pools

021

Fabrication taxonomy

Follow-up worth funding? Y - 022 entropy-signature and 030 confidence-calibration: the gate-blind high-cosine fabrications (entity-sub at cos>0.99) are the exact cases a token-entropy or decoder-confidence detector must catch — this experiment hands them a labeled positive set (out/labels_v1.json + pairs.json). - 026 decoding-strategy: nucleus>greedy head-start confirmed here (do not over-claim; beam not tested). - 029 audit-v2: run the same corpus through decode-parse+NLI to reconcile the 3.6% (semantic) vs the C4

022

Entropy signature

Follow-up worth funding? Y (targeted) - The headline for the program: entropy ≈ cosine, not additive → a fabrication monitor gains little from stacking them; pick entropy when you want a *decode-time, no-re-encode* flag (086 latent-monitor). - 030 confidence-calibration: per-token-class PAV on these exact scores (does p_top1 calibrate to fidelity?) is the natural next step; 022 hands it out/scores.json (per-token stats + labels). - 087 fail-closed-decoder: an entropy/margin abstention threshold at decode time, ROC

023

Off-manifold dose-response

Follow-up worth funding? — Weak Y The clean finding "the decoder does not fail open off-manifold; it fails *closed* into garble except along inter-sentence chords" is a useful safety-relevant map. A worthwhile extension is a sampling-decoder rerun of the chord direction only (where fabrication concentrates) at fine dose spacing around 0.35–0.6, to see whether sampling widens the fluent-wrong shoulder that greedy suppresses. Lower priority than the 024/086 prior-implausibility detector thread.

024

LM-prior mechanism

Follow-up worth funding? Y The "original token is prior-implausible (rank 762)" signature is a candidate fabrication *detector feature* orthogonal to entropy (022) and cosine — worth adding to 086's monitor ensemble and testing on 023's off-manifold fabs (where fabrication is dialled up). Also: does steering z toward the prior's preferred token (025 fabrication-axis) reproduce these substitutions?

028

SAE decomposition

Follow-up worth funding? Y (weak-moderate). (1) Width sweep h∈{8192,16384,32768,65536} (all on box) — does the length/density difficulty axis persist and does m90 scale with width? feeds 044 splitting-absorption / 045 k-atom-curve. (2) Residual probing (043): probe the SAE residual (the lossy 0.45) for length/surface — is the un-reconstructed part exactly the length/rarity signal? (3) Density-controlled fabrication: does the recon↔fabrication link survive length control (partial correlation)?

029

Decode-verify auditor (v2)

Follow-up worth funding? Y (narrow). The NLI/decode-then-verify path is decisively refuted for TAE monitoring — do not pursue sentence-level NLI. The failures are all slot-level (entity/number/role), so a targeted structured-slot checker (extract (agent,verb,patient,numbers,named-entities) from orig and recon, compare slots) is the mechanistically-indicated next auditor — connects to 086 (latent-monitor) and the binding-program's who-did-what focus. The standing lesson stands reinforced: no cheap post-hoc auditor r

030

Confidence calibration

Follow-up worth funding? Y (narrow). Feed the PAV-recalibrated confidence + the graded risk-coverage curve directly into 087 (fail-closed decoder) as the abstention primitive — the honest lever is coverage tradeoff, not a precise fab gate. The under-confidence finding suggests temperature/label-smoothing recalibration of the decoder itself could sharpen p_top1, but the ceiling is low (raw Brier ≈ base rate).

033

Analogy operator battery

Follow-up worth funding? Y (narrow) Voice is the cleanest generative probe yet of the binding gap: fit the passive offset, apply, and measure argument-order fidelity as a scalar. Feeds 034 (negation operator — here 1.00, causally steerable next), 064 (translation-as-offset), 079/081 (operator composition/emergence across rungs), 080 (causal steering along these directions). The diff_align diagnostic is a cheap pre-screen for "is transform T a linear operator" without any decoding.

034

Causal negation operator

Follow-up worth funding? Y (narrow). - The negation direction as a latent monitor / steering primitive (feeds 080 sonar-causal-steering, 086 latent-monitor): a single vector whose addition provably flips a semantic property. - diff_align again pre-screens linearity (0.65 natural vs 0.84 templated predicts the mild degradation) — cheap "is T linear" gauge, consistent with 033. - Dose-optimum >1.0 for natural text (offset under-scaled) is a general lesson for latent operators fit by mean-difference on short templates

035

Operator curvature

Follow-up worth funding? Y (narrow). The operator field is provably curved (a global offset mis-serves length extremes by ~18°) yet application-robust at this scale — so curvature matters for *precision* steering / larger edits, not for coarse transforms. Worth: (1) test whether curvature breaks application at LOWER cosine budgets or for the HALF-operators (voice arg-swap, 033) where headroom exists; (2) a length-conditioned (curved) offset vs global for long-text steering (feeds 080); (3) parallel-transport the of

036

Norm semantics

Follow-up worth funding? Weak-Y (narrow). - The norm↔specificity link is the only positive: worth a clean confirm with proper NE tagging + a controlled rare-word-injection stimulus (does adding a rare proper noun raise ‖z‖ at fixed length?) — connects 066 bits-accounting (norm as a crude bits gauge) and 070/071 numeric/entity fidelity. - The scale-invariance result is a useful negative primitive for steering work (080/082/086): attacks/edits should move *direction*, not magnitude; a magnitude-only perturbation is a

037

Anisotropy audit

Follow-up worth funding? Weak-Y (as a certificate, not a new hunt). The valuable output is the certificate itself: *SONAR z is near-isotropic, so campaign geometry results are not basis-artifacts, and the correct confound to police is norm-scale/raw-Euclidean density (036 lesson), not classic anisotropy.* Narrow extensions: full ZCA-whitening stress test (does inflating tail dims break anything?) and re-running 040 (whitening-robustness) can now cite this μ+PC fit as the shared normalization.

041

Crosscoders across depth

Follow-up worth funding? Weak-Y. The clean negative (depth-partitioned dictionary, near-zero shared atoms) is itself informative and the z-reconstruction-improvement-under-shared-encoding is a genuine lead. A tighter re-run (match the 028 corpus + h/k so z-FVU lands in 0.44–0.55; add a real coherence judge; sweep the share threshold τ) would firm up whether ANY concepts persist across depth or whether SONAR's dictionary is strictly stratified. Not urgent.

042

The residual is structured

Follow-up worth funding? Y (narrow) (1) Row 043 residual-probing is already queued and directly extends this: WHAT information lives in r (probes for length/topic/register/syntax). (2) A properly powered residual dictionary probe (larger h, aux-k, more epochs, power-certified on x first) to settle P1's open question. (3) The 042 finding reframes the 041–050 block: at FVU ~0.7 the SAE dictionary explains a minority of sentence content on generic text — worth one row testing whether campaign-corpus FVU 0.44 vs pile 0

043

What the dictionary drops

Follow-up worth funding? Y The dictionary misses most probe-readable content at k=32 — a k/width-sweep of "probe recovery vs FVU" (does the residual's probe content vanish as FVU→0, or plateau?) would say whether SAE dictionaries systematically privilege some feature types (topic vs length vs lexical) — directly relevant to interp claims that SAE features "explain" the representation.

044

Width buys nothing

Follow-up worth funding? Y (modest) Row 045 (k-scaling at fixed h) is already queued and is now the decisive arm: if FVU/residual content is also k-invariant, the SAE objective itself caps readable content; if k buys it back, the 042/043 residual is "content beyond the top-k budget", which reframes all w40-family dictionaries. Also cheap and worth it: one activation-correlation rerun of the novelty matcher (fixes the gated cell with data already on the box).

045

The k-atom curve

Follow-up worth funding? Y Two forks: (1) TRAINED k-ladder (k ∈ {64,128,256} same recipe, ~min/train at h16384) — does trained k=128 close the decode gap that inference-k cannot, or does FVU-vs-content dissociate there too? (2) The FVU/chrF dissociation itself: what do rank-33-128 atoms add that decoders love but L2 hates? (candidates: low-norm content directions vs high-norm scaffold; connects 036 norm-semantics.)

046

Atoms are meaning-indexed

Follow-up worth funding? Y (narrow). 1. The graded matcher (margin / profile-corr instead of saturated argmax-identity) is now the right instrument for the typology gradient AND for re-running 044's gated novelty cell (activation-correlation across widths/seeds) — data already on the box. 2. The centering-resistant ~19% exclusive core (557-atom union): are these script/tokenizer detectors or content? Cheap: top-activating FLORES sentences per exclusive atom.

047

Frames vs atoms

Follow-up worth funding? Y (one specific). PCA-50 > SAE off-distribution (this corpus AND 043's battery FVU 0.826) suggests the SAE's advantage over cheap global bases is pile-specific — a small "FVU vs corpus" grid (pile/FrameNet/battery/FLORES × {SAE, PCA-k, frame/label bases}) would say whether the dictionary's variance story generalizes at all. The abstract-vs-concrete frame split (b) is also a clean handle on WHAT KIND of meanings atoms are (046 follow-up: content vs topic).

048

Atom ontogeny

Follow-up worth funding? Y (modest) The 4000→26000 gap hides where the last quarter of atoms locks in; the ladder trainer saves F1-crossing milestones, so a cheap rerun of rung A with log-spaced step ckpts (shape B-style dense trajectory, ~70 min GPU) would resolve the tail and test the flicker/"final-only" 25%. Also: the untrained-encoder atom inventory (~16% of ceiling, 0.31 of frequent quartile) is a clean handle on "what SAEs find that isn't learned" — worth cross-referencing with 047's topic-selectivity labels

049

Paraphrase invariance

Follow-up worth funding? Y (narrow). 1. The 40%-semantic / 40%-surface per-atom split: rerun semantic index with PAWS pairs stratified by shared-content-word count, or with role-swap minimal pairs (061/062 battery) to isolate proposition-level atoms from topic atoms — directly tests whether ANY atom encodes structure rather than field. 2. In-domain low-overlap paraphrase cell (backtranslation of the same PAWS sentences) to confirm the "rewording ≡ translation" equivalence without the QQP domain confound.

051

Embedder binding sweep

Follow-up worth funding? Y (narrow). The instrument-failure is itself the finding and it directs the next steps: (1) row 053 role-retrieval prompts to try to lift the within-construction ceiling above 0.9 (restore power) before re-asking the binding question; (2) larger sentence embedders to test whether the weak-role regime is capacity-bound; (3) row 052 (LASER/LaBSE MT objective) as pre-registered. Until a small embedder passes the positive control, "no binding in small embedders" stays uncertified.

052

LASER vs LaBSE: the objective

Follow-up worth funding? Y (narrow). (1) Pooling-swap control (attention-pool LASER's states / max-pool SONAR's) to isolate the decoder-objective vs pooling contribution to the SONAR regime. (2) The LaBSE lexically-anchored possessor cell (0.986) deserves promotion into the battery as a named cell ("content-addressed binding") — it is the first non-SONAR encoder to pass the s↔of positive control. (3) German case-binding is now SONAR-specific: test NLLB-encoder (SONAR's parent) on the 061 battery to locate where it

054

The concept-space LM

Follow-up worth funding? Y (narrow): (i) battery on states at position k>1 under role-relevant multi-sentence contexts (tests planning pressure properly); (ii) role-diagnostic probe pairs with novel subjects to fix the forced-choice instrument; (iii) denoiser tower depths. As a cheap fold-in, not a new flagship row.

056

Cross-attention doesn't bind either

Follow-up worth funding? Y (narrow, fold-in) One cell, not a row: run R-decl-incongruent + QA-objrel on one strong modern reranker (bge-reranker-v2-m3 or an LLM judge-scorer) to test whether the below-chance surface inversion is a MiniLM-capacity artifact or a general property of relevance-trained cross-attention; plus promote the objrel QA inversion into the battery as a standing behavioral cell.

060

The judgment anchor

Follow-up worth funding? Y — the human study itself. STUDY_DESIGN.md is runnable as written (~$1,000, N=200 recruit / 160 analyzed, Prolific + jsPsych). It is the single missing anchor for the whole 051–060 arc: it decides whether "embeddings don't bind roles" is a deviation from human similarity or a faithful model of its casual regime. Everything else in the arc is now bottlenecked on that anchor, not on more encoder cells.

061

Case-marking languages

Follow-up worth funding? Y (narrow, breaker-first). 1. Breaker on German before any promotion: MLP readout; more orders (German OVS variants, scrambled subordinate clauses); does the signal survive a null where case markers are swapped to break der↔den? Is 0.668 an artifact of the article being a separable adjacent token (bag-of-words "den+Noun")? → feeds 062 (crosslingual-role-transfer), 073. 2. Morpheme-separability hypothesis as a first-class claim: predict binding strength from how free-standing the case expone

062

Cross-lingual role transfer

Follow-up worth funding? Y (two narrow threads). 1. Transfer-cell redesign that can extrapolate: score RELATIVE codes instead of focal-canonicalized absolutes — e.g. probe on swap-pair differences z(AB)−z(BA) (label = which of the two orders is agent-first for the SAME noun pair), which is vocab-general by construction; then rerun the crosslingual matrix. This is the fix that lets the row's actual question be answered. 2. Japanese is now the binding thread: 0.696 across two disjoint lexicons is the strongest order-

063

The capacity tax

Follow-up worth funding? Y (narrow). 1. Row 064 (language-vector) is now sharpened: the offset is removal-safe and removal-HELPS near the knee — test the reverse direction (adding μ_L − μ_eng to an English z: does it translate? at what fidelity cost vs token-switch translation?). 2. Script-fair fidelity metric (word/morpheme-level or punctuation-normalized) to de-confound the CJK level tax; cheap re-analysis of the saved decodes. 3. The "centering frees capacity past the knee" effect (+5 chrF for deu/tur/jpn) as a

064

The language vector

Follow-up worth funding? Y (narrow). 1. What IS the 91% residual? It is sentence-specific but averages to ~0: test whether it is (i) encoder noise (re-encode paraphrases), (ii) content-language interaction (probe residual for POS/morphology), or (iii) chrF-invisible register/style. 2. 065 midpoint tie-in: the no-code-switching winner-take-all here says decoder output language is discrete in the token — 065 should test WITHIN-z mixtures (interpolate z_eng↔z_deu of the SAME sentence, fixed token) instead of adding v.

065

Code-switching monolingualized

Follow-up worth funding? Y (narrow). 1. The p2 leak channel: what exactly survives in the 36%-mixed clause-switch decodes (function words? seam-local material?) — the only place SONAR emits mixed text; cheap re-analysis of saved decodes. 2. Encoder-token dose on mixed input: cos(mat,emb) 0.60–0.85 says the source token matters ~10× more for mixed than pure input — measure properly with monolingual wrong-token baselines (one encode pass). 3. Preservation gradient deu>tur>jpn (script distance): test with Latin-script

066

The bits budget

Follow-up worth funding? Y (narrow). 1. Prereg the specific-bits ladder properly: random non-adjacent cross-domain mismatched z, item-level pairing, all 6 langs — freeze the ceiling + knee criteria on THAT curve. 2. "Bits present, decode can't surface": beam/sampling/constrained decode on jpn near the knee (ties 026-decoding-strategy) — can better search recover the chrF the bits say is there? 3. Connect the ~460-bit ceiling to 045's k-atom curve / 028 SAE decomposition: bits-per-atom accounting.

067

The dimension ladder

Follow-up worth funding? Y (cheap). (a) Freeze the absolute-τ knee (or a bits-demand knee) as the primary and add d_z ∈ {32, 128} rungs (~3.4 GPU-h) to localize the saturation point of absolute fidelity between 64 and 256; (b) a d_model-widened control (rank ceiling moved to 512) would separate objective from architecture above 256 — the one question this design provably cannot answer.

068

The budget is in bits

Follow-up worth funding? Y (narrow). 1. Add a REAL high-perplexity NATURAL tier (code, math notation, dense named-entity text, or non-English transliterated) to get natural per-token separation ≥1.5 bits and score P4 cleanly on non-artificial text — the RAND result predicts a leftward natural knee too. 2. Fitted-changepoint knee + bootstrap CIs on the chrF curves (frozen-criterion knee is coarse); regress knee_tok on measured rate across all 4 tiers (slope ≈ −budget/rate²). 3. Codex-judge the fabrication-vs-omissio

069

Conjunction is subadditive

Follow-up worth funding? Y (narrow). 1. Relatedness as a CONTINUUM: regress per-pair R on chrF(A,B)/embedding-sim across many intermediate levels — is there a sharp redundancy threshold or smooth exchange rate between shared bits and stored bits? 2. Capacity-gated position effect dose-response: sweep conj length through the knee and watch the first-clause advantage switch on (ties 066d/068; clean causal handle on front-loading). 3. Operator cost: swap ", and" for "but/or/;/because" (ties 033 composition operators)

070

Numbers don't cliff

Follow-up worth funding? Y (narrow). 1. Length dose-response through the knee at fixed D: sweep context 32→144 tok and watch the D=8/12 survival curve collapse — locate the number-fidelity knee vs the chrF knee (ties 068/069 capacity-gating; is digit loss the *first* thing to go as bits run out?). 2. Digit-position error profile: the leading-digit-preserving / tail-corrupting pattern suggests a positional fidelity gradient inside the number — measure per-digit-position exact rate (are units-place digits lost first?

071

The entity ceiling

Follow-up worth funding? Y (narrow, not a new row) (1) push L* higher (80/120 tok) to trace how the matched count-ceiling scales with budget — does ~3 rise linearly with bits like 070's I_spec? (2) matched cells with entity sentence at FIXED position + filler varied, to cleanly separate length-load from entity-count; (3) mixed-rarity items (famous + rare in one z) to re-test rarest-first with real rarity variance (here all ~12 bits, low variance → the null is under-powered); (4) merge-mechanism probe — are two→one

072

What gets deleted first

Follow-up worth funding? Y (narrow) (1) Separate object-role from medial-position: put the patient sentence-initial (passive / OSV templates) and re-measure — is it the *object relation* or the *buried position* that dooms it? (2) Formal codex neutral judge (~100) on ACTION-paraphrase and PATIENT-merge calls to tighten the lower bounds. (3) Push the same 6-role ladder with a bit-budget x-axis (not K) to place each role's half-survival demand on the common 460-bit scale. (4) Object-collapse mechanism: is the repeate

073

When the order code forms

Follow-up worth funding? Y (modest) A cheap retrain of A_D (the exact anti-transfer organism) with log-spaced milestone ckpts through 11%→100% (~1 GPU-h) would (a) remove the A_P90 arm caveat and (b) resolve whether the mid-training rise has any fine abrupt structure hidden in A_D's coarse tail. The locked surfX↔roleXflip co-movement is a clean handle for a causal test: ablate the linear surface-position direction at each milestone and check whether role within-ceiling survives (does role readout depend on the posi

075

Anatomy up to rotation

Follow-up worth funding? Y (modest) Two cheap extensions would sharpen this into a promotable claim: (1) repeat the seed atlas at one higher rung / larger d_z (~6 GPU-h) to test whether "stable up to rotation" is a property of this bottleneck or of ladder-TAEs generally — if the aligned-cosine recovery and the invariant list survive a capacity change, the anatomy claim strengthens sharply; (2) the number-operator anomaly (diff_align 0.35, aligned recovery only 0.59) is a clean handle on "which linguistic operators

076

Curriculum splits the anatomy

Follow-up worth funding? Y-modest. The reconstruction path-dependence is a clean, mechanistic recency/forgetting result that a 1-GPU-hour follow-up could sharpen: (a) mid-vs-final checkpoint anatomy for a curriculum arm (does the *relational* code peak-and-hold while reconstruction peaks-and-forgets? the best-f1 milestones suggest yes); (b) a replay/interleave control to confirm the effect is terminal-recency, not curriculum per se; (c) does a length curriculum with a *balanced* final phase recover val_f1 to the 0.

077

Objective hysteresis

Follow-up worth funding? YES (narrow). The surface-code first-objective imprint (surfX m>1, ‖z‖ m≈0.7) is the surprising, block-I-enriching result and deserves a 2-seed × 2-switch-fraction confirmation plus a layer/step trace of WHEN the surface code freezes (ties to 073's "forms mid-training"). The recency-of-reconstruction result is already well-supported (076+077) and does not need more runs.

080

Steering is causal

Follow-up worth funding? Y — natural-sentence dose-response (where the over-steer break should reappear, per 034) + push α past 2 to localize the fluency break-point + steer a CLOSED-CLASS marker on natural text. The causal claim is now clean at T3; a natural-text replication with the codex fluency gate would be the promotion path above T3.

081

When operators install

Follow-up worth funding? Y (modest). The sharp result — causal installability is decoder-gated and tense never installs — invites: (1) does a stronger decoder (SONAR-scale ladder or more dec_layers) rescue tense's causal steerability, or is tense's linear offset fundamentally non-causal (probing the 078 binding wall from the time axis)? (2) a finer log-spaced A_D milestone retrain to pin the decoder-onset step exactly and test whether negation/vertical install at the SAME decoder-competence threshold (ceiling≈0.5)

082

The cosine gate's blind spot

Follow-up worth funding? Y (narrow): (1) codex neutral-judge cross-check to promote above T3; (2) natural (non-templated) sentences — does the negation/entity cost ordering hold off-manifold?; (3) a *logical-flip-aware* auxiliary probe (negation/quantifier direction) as the missing orthogonal check the cosine gate needs.

083

Covert capacity, and its catch

Follow-up worth funding? Y (narrow): (1) load the real w40 SAE and test the SAE-residual carrier specifically (does dictionary structure raise or lower detectability?); (2) natural (non-templated) sentences — does the ~100-bit capacity and the nn-cos AUC hold off-manifold?; (3) an *adaptive* payload that minimises off-manifold nn-cos deviation (project onto the data manifold's tangent) — the true worst-case for a manifold monitor, to size the residual undetectable capacity.

084

Canaries survive retrieval

Follow-up worth funding? Y (narrow, feeds 085). The clean positive here sharpens 085: role-swap alone does not break retrieval, so 085 should test whether a distractor that is topic/surface-closer to the query BUT role-wrong can beat the role-correct canary — i.e. can surface similarity overpower the +0.16 role margin? Also: heavier/cross-lingual paraphrase queries and a natural-text canary corpus to see where the role margin finally collapses.

085

One word breaks RAG

Follow-up worth funding? Y (narrow). (1) Re-run with 084's tightly-matched paraphrase query as the incumbent to report the flip threshold against the full +0.159 role margin directly (predict ≈2 words). (2) Natural-corpus base rate: how often do real passages carry a role-wrong sentence with incidental surface overlap to the query? (3) The defensive fix: a role/agent-consistency check between query subject and retrieved-passage agent as the orthogonal guard the cosine retriever lacks (ties to 082's "cosine needs a

086

The ensemble monitor

Follow-up worth funding? Y (narrow): (1) tangent-adaptive covert payload (083 follow-up) vs the decode- consistency check — size the truly-undetectable capacity; (2) natural (non-templated) sentences — does the re-encode floor loosen enough that cos_ref stops dominating the primary?; (3) a learned (not prototype) logical probe to lift negation AUC above 0.85 in the paraphrase-admitting regime.

087

Fail-closed decoding

Follow-up worth funding? YES (small): pair the density gate with a round-trip-consistency gate and re-run the tradeoff — hypothesis is that density ∧ round-trip closes both the off-manifold AND the interpolation doors while still keeping clean+covert. - Caveats: equal 5-category test mix (population fab is a design choice; the density-ordering result is mix-robust but the absolute no-abstention rate is not). Simple templated clean stimuli (low clean-fab baseline). Single seed, single embedder (SONAR). nn-cos densit

088

Membership is memorization

Follow-up worth funding? Y (narrow). 1. Memorization dose-response: sweep train-set size / epochs from the STD (0.55) to the MEM (1.0) regime to map AUC vs the 066 recon-fidelity gap — quantify "how much memorization ⇒ how much leakage." Cheap (forward-only, reuse this harness). 2. Two-bank density to test the manifold-density membership hypothesis (031/083) without the bank-distribution artifact. 3. SONAR proper needs a *real* member/non-member corpus (e.g. a known NLLB bitext slice vs a post-cutoff/held-out slice

090

The probe-power kit

Follow-up worth funding? Y (as infrastructure, not a study). The kit is now the standing pre-flight for every future probe-vs-representation row: call ppk.certify(Z, y, target_effect, readout, groups) and only report a null if it PASSes. Natural extensions (cheap): (1) an mlp-readout certificate pass for reps with nonlinear-only signal; (2) a helper that maps a real intrinsic AUC back to an equivalent Cohen's d so callers can state "our null rules out effects ≥ d*"; (3) wire it as an assertion into the battery's ve

091

The canonicalization checker

Follow-up worth funding? Y (narrow) Wire canon_check.certify() as a pre-flight gate into the binding battery itself (load_task could assert it), and extend the item-schema adapter to the ladder/SONAR retained-state probe rows so every future binding probe self-certifies its labels.

092

The norm linter

Follow-up worth funding? Weak-Y (narrow, cheap). - Re-encode ~40 real 033/078 operator before/after pairs (GPU, seconds) → run the linter on a real operator delta as the PASS anchor (closes the one synthetic gap). - Wire norm_linter.lint_delta as a pre-flight assert into future offset/operator rows (034/046/062/063/064/080) — it is a 2-second import norm_linter call.

093

Adversarial stimuli evolution

Follow-up worth funding? Y (narrow). Give codex the filler-pool + lexical-diversity levers (the 006 axis) as the attack surface and re-run the loop with the full battery + CIs: that is the one place a plausible+well-formed z_bag/STIMULI_INVALID break might actually exist. Absent that, the template-level conclusion (probe + canon_check gate robust to adversarial construction evolution) is a solid T3 instrument-validation result.

094

The prereg engine

Follow-up worth funding? (Y, narrow) Y — wire prereg_engine into the launch/harvest of future rows so the PREREG_LITE emits a PREREG.json (frozen certificate) at launch and RESULT.md's Brier table is generated by Verdict.brier_table() at harvest. Cheap, and it makes the "score exactly as frozen" rule mechanical rather than manual. Sibling of the 090/091/092 kits; block-L instrument suite now covers probe-power, canonicalization, norm-linting, and (this row) the prereg/Brier method itself.

095

The Bayesian ledger

Follow-up worth funding? (Y — narrow) Y: (1) wire the ledger into harvest so each new row auto-appends its Evidence (row + Brier + one-line direction) and the posteriors update mechanically — the campaign's live scoreboard. (2) Replace analyst strength labels with a rule that reads effect size + gate outcome directly from out/results.json (removes the one subjective step). (3) A proper independence discount (down-weight correlated rows within a block) would let C1 report an honest sub-clamp posterior.

096

Why bag codes win

Follow-up worth funding? Yes, conditionally. Corrected, falsifiable prediction for the real system: 006's transition location is governed by the strength of the residual order/position signal that survives pooling (β in the toy), and is roughly insensitive to both training-token count and bottleneck width — the opposite of a naive capacity story. Discriminating experiments: (i) a data sweep at fixed capacity should NOT move 006's knee (toy: T-independent); (ii) a bottleneck-width sweep (ties to 067) should also NOT

097

What survives pooling

Follow-up worth funding? Yes, narrow. Cheap, sharp real-SONAR test of the two robust claims: (i) extract encoder contextual states (011–020 rows already did this), mean-pool, and show order recovery is per-instance yes / global linear axis no (the cross-pair AUC≈chance vs within-bag high contrast) — a direct check of the content-entanglement mechanism; (ii) confirm the surviving-order subspace is low-rank (top-few PCs of z−bag over shuffles capture the order variance) and small-amplitude, tying its strength to 096'

098

The knee from first principles

Follow-up worth funding? Y (narrow, cheap). 1. Compute δ(D) explicitly: derive the recall-0.85 and chrF-0.8× budgets from a distortion functional on the 066/068 I_spec curve, turning the empirical 0.41 / 0.78 into a predicted R(D). 2. Re-decode the anatomy capacity cell to get its true sentence length distribution (closes the P6 soft joint) and a full v2 CI/bimodality flag on 15.79 (still owed per the ladder prereg). 3. Regress knee_tok on measured rate across 068's four tiers (slope ≈ −C_D/r²) as a direct R(D) tes

099

The z-explorer

Follow-up worth funding? Y (low cost) — this is the campaign's public face; worth (a) precomputing a live-ish "type your own sentence" path if an in-browser SONAR shim ever exists, and (b) adding the 034 negation-operator and 083 stego demos. But as-is it already covers the five headline phenomena end-to-end.

100

TAE-Bench release

Follow-up worth funding? Y — (i) refactor the battery's per-cell readout to release memory (process isolation or explicit del + gc), enabling the full SONAR sweep; (ii) extend the --encoder hook to more models; (iii) promote the flagship null past T3 with an independent breaker pass.