Skip to content
Shadow Theory

Paper 3 · Appendix CLearning

Exact examples that discipline the inference

A genuine refinement can worsen a different prediction objective; a valid safe replacement can be limited to its measured panel; an event can refute a whole candidate bank. The proofs show precisely why.

Section 14 of 18

C Analytical controls on interpretation

The following finite arguments support specific diagnoses. They are not convergence theorems for the implemented learners and do not replace the interface results in the companion theory paper.

C.1 A true refinement can worsen the selected worst-context error

Take four equally weighted histories a,b,c,da,b,c,d with deterministic next-output probabilities (0,1,0,1)(0,1,0,1). A one-state maximum-likelihood fit is q=1/2q=1/2, with worst-history TV 1/21/2. A valid counterexample requires aa and bb to be separated. The refined partition {b},{a,c,d}\{b\},\{a,c,d\} fits probabilities 11 and 1/31/3. It obeys that constraint and improves average negative log likelihood from log⁡2\log2 to

−2log⁡(2/3)−log⁡(1/3)4<log⁡2.\frac{-2\log(2/3)-\log(1/3)}{4}<\log2. (21)

At history dd, however, TV becomes 2/32/3. The statement concerns the result of likelihood refitting, not the optimum under a fixed worst-case loss. If H0⊆H1\HH_0\subseteq\HH_1 and the same exact risk RR is minimized globally, then inf⁡H1R≤inf⁡H0R\inf_{\HH_1}R\leq\inf_{\HH_0}R is immediate. A richer class does not itself worsen the best available minimax risk.

This counterexample was analyzed after OII-2 scoring. It motivated OII-3; it was not a preregistered prediction proving that objective mismatch caused every OII-2 failure.

C.2 What the safe-replacement rule certifies

Let RE0(M)R_{\E_0}(M) be worst-context TV on one fixed measured panel E0\E_0. Suppose a simultaneous confidence event gives L(M)≤RE0(M)≤U(M)L(M)\leq R_{\E_0}(M)\leq U(M) for every compared candidate, including the incumbent M0M_0. Then

U(M1)+η≤L(M0)⟹RE0(M1)+η≤RE0(M0).U(M_1)+\eta\leq L(M_0) \quad\Longrightarrow\quad R_{\E_0}(M_1)+\eta\leq R_{\E_0}(M_0). (22)

The proof is the chain R(M1)+η≤U(M1)+η≤L(M0)≤R(M0)R(M_1)+\eta\leq U(M_1)+\eta\leq L(M_0)\leq R(M_0). Comparing two upper bounds alone does not give this implication. Uniformity must cover the selected candidates; using adaptively chosen models with merely pointwise intervals would require another justification.

OII-3 obtained a candidate-uniform bound by bounding the true record law. For a pilot-selected atom set A0A_0, independent calibration supplied simultaneous lower mass bounds lhl_h. Every candidate QQ then satisfies

TV⁡(P,Q)=1−∑hmin⁡{P(h),Q(h)}≤1−∑h∈A0min⁡{lh,Q(h)}. \TV(P,Q)=1-\sum_h\min\{P(h),Q(h)\} \leq 1-\sum_{h\in A_0}\min\{l_h,Q(h)\}. (23)

The omitted mass remains uncertain rather than being declared impossible. Independent resets, stationarity and the stated multiple-comparison allocation are hypotheses of the probability guarantee. Whether the bound is useful depends on its resolution. It does not constrain contexts outside E0\E_0 without a coverage or structural argument.

For a convex bank with vertex laws QjQ_j, suppose a convex upper objective UU has value cc on every vertex and a valid lower certificate proves inf⁡wU(∑jwjQj)≥c\inf_wU(\sum_jw_jQ_j)\geq c. Convexity gives the reverse inequality U(∑jwjQj)≤∑jwjU(Qj)=cU(\sum_jw_jQ_j)\leq\sum_jw_jU(Q_j)=c. Thus the objective is identically cc on that bank. This is the conditional reasoning behind the 74 post-reveal flat-objective diagnoses. It does not make all uncertainty-aware objectives flat.

C.3 Exact operational twins and a finite certificate's reach

Two finite generators with matched initial laws and identical admitted complete record laws cannot be separated by a test using only that contract. Hidden coordinates can be split, relabeled or added without changing those laws. A learner's finite nondetection is weaker: two distinct laws can look identical on a small observation sample. Conversely, state labels from two numerical fits can differ even when their induced public laws agree.

An event witness of Equation 9 is sufficient to exclude the entire frozen convex bank. Failure to find one does not establish that any member is adequate. Similarly, the common-history inequality in Equation 11 is a sufficient refutation of a proposed shared prediction, while Equation 17 shows that it is not a complete criterion for a whole merged class.