# S2 — a polarity/count premium, but negative role discrimination Frozen W40 invariant masks outperform historical-activity-matched noninvariant masks on the pooled controlled task, but the positive premium is driven by polarity and count. Role-reversal discrimination is worse than the matched masks and belowchance. These results do not establish general proposition-semantic enrichment or monosemantic atoms. | Inference | Family | Invariant rankutility | Matched masks | Difference | |---|---|---:|---:|---:| | per_sample | role | 0.1909 | 0.2843 | -0.0934 | | per_sample | polarity | 0.9865 | 0.5867 | +0.3998 | | per_sample | count | 1.0000 | 0.8416 | +0.1584 | | threshold | role | 0.2533 | 0.3372 | -0.0840 | | threshold | polarity | 0.9906 | 0.5804 | +0.4102 | | threshold | count | 1.0000 | 0.8808 | +0.1192 | The equal-weight pooled premium is+0.1549 [0.1439,0.1657] forper-sampleTopK and+0.1484 [0.1354,0.1615] forfixedthresholds. These are conditional16-lexical-block bootstrap intervals acrossfixedmodels/banks, not population ortraining uncertainty. Bothvalid support covers99.6%/99.5% of invariant/control comparisons; an exact descriptive decomposition attributes only about0.0017/0.0011 of those premiums to support-asymmetry terms. Most premium therefore comes from comparisons where bothcodes are defined, not the0.5abstention convention. The historical-only matching-overlap sensitivity retains a positiveprimary premium ofabout0.115 forbothrules, but its harder paraphrase-versus-same-surface-counterfactual premium shrinks to0.0174 (per-sample interval crosseszero) and0.0276 (threshold). Trimming changes the featureestimand; retainedcount/mass andbalance are saved. Approximate matching does not eliminate all selectionconfounding. Nativefullz andc_pool cosine eachscore0 onroles and1 onpolarity/count. Thus undoing the fitted positional rotation beforepooling did not make this cosine comparison role-invariant. Fulltextbag scores0.5/0.75/1 respectively (macro0.75); it has originaltext access and illustrates lexical solutions, not an access-matched latent baseline. Invariant macros are0.7258/0.7479. A positivepremium against matchedmasks is not a victory over everybaseline. All96count clusters score1 forinvariantmasks;32 remain under the prespecified equal-actual-token-length restriction, stillscore1. The task can be solved by literal number/polarity cues and does not establish language-independent abstractcount semantics. Role reversal changes proposition identity while allowing bothdirections to be true. Individual entity/predicate features can legitimately survive it; poor cosine discrimination does not prove role information absent fromthe codes. Data:288clusters,16lexicalblocks,1152unique strings crossingmeaning andsurface, withbothdirections/forms anchored. Independent source36 review andClaude's unlabeledsource36 review passed; unusual occupation/event combinations andreferent assumptions remain explicit. C_pool compatibility andnativepool reconstruction gates passed; no new maskselection/calibration used thisbank. All8historicalcheckpoints andbothrules were reported. P1–P3 occurred;P4 (c_poolroleutility exceedingnativez) failed. Source/protocol/code/assets/controlmasks were frozen at02:45UTC. Independent audit reproduced all371cell aggregate/family metrics,8premiums, andrawsparse cosine rankings forfourrepresentativecells (reviews/S2_AGGREGATE_INDEPENDENT_CHECK.json andS2_RAW_METRIC_INDEPENDENT_CHECK.json). Fullcodes, vectors, tokenlengths, masks, support, scores andpredictions are retained.