Skip to content
Shadow Theory

Paper 3 · Section 5Learning

Three different achievements

All 48 targets fit the chosen general architecture; only 39 receive a passing produced candidate and 37 receive a passing selected one. Exact embeddings and event witnesses explain what those gaps establish.

Section 6 of 18

5 Capacity, produced candidates and selection

5.1 A post-reveal representational-capacity sanity check

An unsuccessful fitted model does not reveal by itself whether the architecture lacks capacity. OII-4 permits a direct post-reveal check. Every source generator has at most 16 hidden states and uses the same action/output form as Equation 1. For a target with n≤16n\leq16, embed its states in the first nn coordinates of a 16-state model. Copy every rational joint transition row and each preparation distribution there; set unused states to emit symbol zero and self-loop, with zero initial mass and no incoming mass from the active subspace. The resulting stochastic instrument reproduces every source row on the active subspace and hence every admitted adaptive finite record law.

The saved construction was checked at 3,072 embedded instrument rows and 192 preparation rows. This elementary embedding is an exact capacity sanity check for the general joint-instrument architecture, not a novel representation theorem. It is not a theorem about a more restrictive HMM that factorizes emission and transition in a prescribed way. It also does not prove that the specific positive-pseudocount, single-initialization fitting procedure can attain the exact boundary parameters from its finite data.

These exact source-derived parameters were constructed only after reveal and were never included as known-correct training candidates. A strictly positive parameterization can approximate the embedded model by mixing each initial distribution and joint row with a sufficiently small uniform component. The imported coupling bound gives eight-call error at most 1−(1−η)91-(1-\eta)^9 when each such change is at most η\eta. This is an approximation-capacity statement, not successful identification by the implemented optimizer.

Thus the remaining failures cannot all be explained by insufficient latent-state cardinality. Finite observations, initialization, regularization, local fitting, candidate construction and selection remain possible contributors. Nor does the embedding imply 16 public predictive states: Equation 8 can attain many distinct posterior values even with a small hidden carrier.

5.2 Certificates against a produced bank

Let Bf={Q1,…,QJ}\BB_f=\{Q_1,\ldots,Q_J\} be the finite bank actually produced on fixture ff. An episode-level mixture has laws Qwe=∑jwjQjeQ_w^e=\sum_jw_jQ_j^e, with the same weights across the experiment family. For any legal event AA in an experiment ee,

TV⁡(Pe,Qwe) ≥ Pe(A)−Qwe(A) ≥ Pe(A)−max⁡jQje(A).\TV(P^e,Q_w^e)\ \geq\ P^e(A)-Q_w^e(A) \ \geq\ P^e(A)-\max_jQ_j^e(A). (9)

If the final quantity exceeds τ\tau, no mixture in that bank passes this test. This elementary certificate is stronger than observing that one optimizer chose poor weights. It does not exclude an unproduced parameterization in the same mathematical architecture. The witness and probability evaluation are post-reveal diagnostics, not queries available to the opaque learner during selection.

Table 7. OII-4 post-reveal candidate-bank diagnostics. “Adequate” in this table means passing all fixed core tests at TV ≤0.15\leq0.15.

BankPassing single
candidate /48
Selected
all-core /48
Selection
misses
Convex-bank
obstructions /48
R393728
S393728
E2828019
G1515033
B3736110
P25241–
O2523222

P selects three trees; O mixes four. The obstruction count 22 refers to the four-tree bank. No P-specific obstruction count is inserted. A missing adequate individual candidate does not exclude an adequate mixture.

5.3 What the counts distinguish

For both R and S, 39 of 48 fixtures have at least one produced candidate passing every core slot, while the selected predictor does so on 37. The two differences are genuine core-panel selection misses in the post-reveal ranking. They do not establish that the selector could have known the oracle ranking from the available calibration sample. Table 8 identifies both cases. The selected model has lower observed calibration NLL in each; the failure is not a tie-breaking error or evidence that the selector saw the hidden core ranking. The alternative rows are post-reveal comparisons of candidates that had already been fitted.

Table 8. The two R/S core-panel selection misses. Lower calibration NLL wins under the original selector; a core-worst TV at most 0.15 passes.

Fixture familyModelRoleCalibration NLLCore worst TV
Leaky occupancyHMM-8Selected0.0832320.152781
Leaky occupancyHMM-16Passing alternative0.0877840.131943
CoinIOALERGIA-1.5Selected0.0864550.196476
CoinHMM-2Passing alternative0.0866540.091927

The rows compare already-produced candidates, ranked after reveal for diagnosis. R and S make the same selections. The coin row shows the lowest-calibration-NLL passing alternative; two additional hidden-state candidates also pass. Full fixture IDs and values are in the audit supplement.

Of the remaining nine fixtures with no adequate individual R/S candidate, eight have a Equation 9 certificate excluding every convex mixture in the bank. The last case remains unresolved at mixture level: failure of every individual member does not imply failure of every mixture. Consequently, the 11 selected-model failures separate into two selection misses, eight certified produced-bank obstructions, and one case lacking an adequate individual candidate but not excluded by this mixture audit.

The tree bank has 22 such obstructions on the same final cohort, versus eight for R/S. That is a reduction in produced-bank insufficiency, not proof that one full architecture is globally more expressive than the other. The capacity construction already shows why those categories must be distinguished.

Original-budget diagnostic progression for 48 fixtures: exact parameterization exists for all 48, an individually passing candidate is produced for 39, and the selected model passes all core tests for 37. The eleven selected failures comprise two selection misses, eight banks excluded by an event witness, and one unresolved mixture case.
Figure 4. Original-budget OII-4 R/S diagnostic separation. The arrows show different questions, not a proof that the learner knew any oracle fact. All counts concern the same 48 fixtures and registered core tests; the separate post-hoc expansion does not overwrite them.

5.4 Diagnosis without forced single causes

There is a hierarchy of questions, not a compulsory single-label taxonomy. An architectural witness answers an existence question. A bank witness excludes a finite collection and its mixtures. A selector miss compares a chosen candidate with other candidates that were produced. None alone identifies the statistical cause of a failed parameter fit.

The frozen study used one start and at most 60 EM updates per size. The post-hoc analysis in Section 6 now directly tests a larger budget and finds five selected repairs among its 11 failures. This is evidence of finite-budget sensitivity, not a diagnosis that every residual fit is a local minimum. Insufficient branch evidence, the fixed pseudocount and selection on limited calibration remain possible contributors. Likewise, source models with identical public laws are non-identifiable in their hidden anatomy; that is not the same failure as predicting the public laws incorrectly.

Table 9. Failure categories and the evidence needed to distinguish them. Categories may coexist; the table is not a mutually exclusive labeling of all failed fixtures.

CategoryWhat this record establishesWhat remains unestablished
Architecture capacityExact post-reveal 16-state embeddings for all final targetsIdentifiability or convergence of the opaque fitter
Produced-bank insufficiencyEight R/S bank witnesses exclude every produced mixtureInadequacy of all parameterizations or other candidate generators
Parameter estimationFitted laws disagree on some source contexts despite architectural capacityA unique attribution to sampling rather than optimization or regularization
Optimization/local fittingFive selected repairs among 11 targeted failures under the post-hoc expanded budgetGlobal optimization, a unique causal decomposition, or repair of every remaining miss
Model selectionTwo R/S candidates were available that passed all core tests but were not selectedThat the available calibration observations identify the oracle winner
Context/data coverageDuplicate and calibration-matching specifications; observed panel-to-core failuresA full unseen-policy or all-future generalization guarantee
Finite-budget uncertaintyOII-3 confirmation limits and flat confidence objectivesAdequacy from finite nondetection or exact equality from overlapping intervals
Operational non-identifiabilityExact public-law twins have different hidden implementationsPermission to excuse wrong predictions of operationally distinct laws