Skip to content
Shadow Theory

Paper 3 · Appendix BLearning

Every stage, with its own denominator

Complete OII-1 through OII-4 tables retain every method, mechanism family and registered weight. The deduplication sensitivities are reported alongside the original comparisons.

Section 13 of 18

B Stage-specific numerical results

The tables below preserve their original study-specific denominators and primary weights. They are not pooled estimates of one population or a ranking across cohorts. OII-1–3 mean errors use their saved conservative score convention; the OII-4 endpoint uses its saved full-law numerical scores with separate event certificates. Dashes mean that the indicated control was not separately evaluated, not that its count was zero.

B.1 OII-1 and OII-2

Table 16. OII-1 common core. The separate 705-slot expanded panel is excluded.

Acquisition/predictorPasses /420Mean TV upperAll-core /28Red /28
active/factorized2300.3690429–
active/memoryless1020.7091503–
active/selected2810.2238911215
passive/factorized2160.37178710–
passive/memoryless1030.7132423–
passive/selected2660.2187081313
Table 17. OII-2 common core and same-data fit ablation. A dash denotes a comparison not separately executed.

MethodPasses /720Mean TV upperAll-core /36Red /36Unsafe
/4,990
refined_active6330.086537211443
refined_passive6300.078505211143
oii1_active4420.263312142185
no_confirmed_constraints6390.08078421––
memoryless1390.7626373––
factorized5020.25643115––

The OII-1 common core contains 420 slots per predictor. The expanded 705-slot panel included learner-selected probes and is not substituted for that common comparison. The two OII-1 unsafe-merge searches visited different sets of histories, so their 56 and 131 witnesses are not directly comparable population rates.

OII-2's no-confirmed-constraint ablation uses the same active dataset, including observations obtained during conditional counterexample replay. It tests the added fitting constraints, not the value of those replay observations themselves. Its red and common-merge outcomes were not separately searched. OII-2's memoryless and factorized controls are retained in their documented role; they should not be confused with the separate OII-1 controls.

B.2 All OII-3 V2 selectors

The labels identify the acquisition lane and selector. “CE” is the conditional-counterexample lane, “disagreement” omits that replay, and “passive” uses passive acquisition. INC is the likelihood-selected incumbent. L is a likelihood mixture; M is the empirical minimax selector; U optimizes a confidence upper bound; SAFE applies the measured-panel replacement rule. BASE, FREE and CON retain their saved candidate-bank/constraint meanings. The OII-2 controls have the same total call allowance but their own allocation rather than the new lanes' dedicated calibration design.

Table 18. All 32 OII-3 V2 selectors on their common 1,008-slot panel. This is not a pooled comparison with the other studies.

Lane/selectorPassesMean TV upperMean fixture-
worst
All-core /42Red /42
ce/INC8430.1046250.3768732219
ce/L_CON8360.1091670.3734162220
ce/L_FREE8360.1091670.3734162220
ce/M_BASE8060.1211510.3684162219
ce/M_CON8060.1211510.3684162219
ce/M_FREE8060.1211510.3684162219
ce/SAFE_CON8390.1061480.3776222219
ce/SAFE_FREE8390.1061480.3776222219
ce/U_CON7140.1734060.4438732021
ce/U_FREE7120.1739710.4438292021
disagreement/INC8490.1108600.4406861922
disagreement/L_CON8470.1102190.4280751922
disagreement/L_FREE8470.1102190.4280751922
disagreement/M_BASE8170.1213040.4146991922
disagreement/M_CON8170.1213040.4146991922
disagreement/M_FREE8170.1213040.4146991922
disagreement/SAFE_CON8260.1196120.4279411922
disagreement/SAFE_FREE8260.1196120.4279411922
disagreement/U_CON7050.1803900.5066861725
disagreement/U_FREE7050.1803900.5066861725
oii2_active8620.0927660.3996142219
oii2_passive8710.0858510.3296472218
passive/INC8300.1187630.3807152219
passive/L_CON8190.1126170.3719722218
passive/L_FREE8190.1126170.3719722218
passive/M_BASE8190.1180670.3743592219
passive/M_CON8190.1180670.3743592219
passive/M_FREE8190.1180670.3743592219
passive/SAFE_CON8300.1187630.3807152219
passive/SAFE_FREE8300.1187630.3807152219
passive/U_CON7370.1656070.4288152119
passive/U_FREE7370.1656070.4288152119

B.3 All OII-4 mechanism families

Table 19. Complete OII-4 family pass counts on the registered primary slots.

FamilyTestsRSBPO
alternating channel484343432413
budget483939394848
challenge482224824
checksum protocol484242434531
coin727070727272
gated queue484848433118
hidden belief484848481616
joint727272726262
leaky occupancy484747474433
lock484646464646
mod counter484848483529
mode register484848402322
parity727272727272
probe484848483322
queue484848484848
receiver483232324848
redundant727272727272
refractory server484848482413
sharedbudget484646464546
stochastic gate484848484242
transaction buffer484848484444
weak484848484848

B.4 Deduplication sensitivity

Table 20. Two retrospective deduplication sensitivities; neither replaces the 1,152-slot primary analysis.

MethodLiteral passes /1,140Behavior passes /1,135All-core /48
R1052104837
S1052104837
E90289828
G55655315
B1042103836
P95895424
O85885423
H110

Finite behavior means equality for all possible output strings at the matched horizon, not equality only on paths of one particular source. Calibration-matching contexts remain included. All-core counts are unchanged.

Removing duplicate within-fixture specifications or finite policy behaviors changes the primary weighting. This table is therefore a retrospective sensitivity only. It retains specification matches with calibration; deduplication is not a leave-calibration-out test. All-core fixture pass counts are unchanged because removal of an identical duplicated test cannot alter whether every distinct core context passed. The reported decimal law scores are those already stored; no models were refitted for the sensitivity.