Paper 2 · Section 4Boundary
Recover the full experiment from a partial view
Directed statistical deficiency measures the information lost when an observation is deleted. Exact formulas, optimal decoders and a decision witness resolve the masked family across its full parameter interval.
4 Robust reconstruction of joint experiments
4.1 The information lost by deleting a view
The obstruction concerns exact target selection, not the existence of stable measures of joint information. A different question is available: after retaining only part of an observation, can one common reconstruction procedure reproduce the full experiment for every unknown preparation?
Let be a nonempty finite preparation set. An experiment has laws on a finite joint alphabet . A retained view is the deterministic projection , with image alphabet and marginal laws .
The loss of reconstructing from is
The stochastic row kernel is the same for all unknown . It may use a declared public context, but not the unknown source label.
This is directed statistical deficiency in a target-first, half- convention, not a new information measure. Statistical comparison has its classical Blackwell/Le Cam lineage [6, 16]; deficiency has also been applied directly to feature quality [24]. Argument order, row versus column kernels, and the factor of two in variational divergence must be translated when comparing formulas.
The decoder reproduces laws. In particular it may resample even the retained coordinates. It need not recover the same realized event, preserve a particular archive carrier, respect a deadline, or be available as a native physical operation. Those stronger requirements belong to the execution contract in Section 5, Section 6. They are not consequences of Equation 4.1.
The minimum in Equation 4.1 exists, lies in , and is a rational linear-program value for rational finite data. For any ,
For experiments on the same parameter set,
More informative retained observations cannot increase reconstruction loss, nor can postprocessing the target. If , with the same projection and preparation labels, then
All these statements preserve the nominated experiment and decoder permissions.
Finite stochastic matrices form a compact polytope and the objective is continuous. The absolute-value epigraph gives a finite linear program with rational coefficients; a rational optimum can be chosen. Insert and between the two target laws and use TV contraction under to obtain Equation 4.2. Composing two reconstruction kernels proves Equation 4.3; the same argument gives the monotonicities. For a fixed , replace both target and input laws by their hatted versions. The first replacement costs at most , and the second at most after contraction. Take infima in both directions to obtain Equation 4.4.
□For a decision loss valued in , any decision procedure based on can be emulated from , using an optimal followed by that procedure, with additional risk at most for every preparation and hence every prior. This forward implication follows from the TV bound on bounded expectations. No full statistical-comparison converse is needed below.
One may record the entire table . A derived deletion diagnostic is
where the empty view is a one-point alphabet. Positive means that each one-coordinate deletion loses experiment information. It is not an additive information decomposition, a physical edge attribution, or an exhaustive notion of integration. In particular, it should not be identified with a partial-information-decomposition atom; unique-information constructions answer their own specified comparison questions [5].
4.2 Genuine informational padding
For a deterministic view , if and only if there is one conditional kernel , supported on , such that
In particular, if
with the same conditional kernel in every preparation, then
for any common retained experiment .
A conditional reconstruction immediately proves the forward zero-loss claim. Conversely, an optimal zero-loss kernel reproduces the full experiment, though it need not initially preserve the view. Put a strictly positive prior on the finite . The reconstructed Markov chain has the original law at its endpoints. Data processing gives . Since is originally a deterministic function of , the reverse inequality holds. Thus in the original experiment. Its conditional law supplies , common to every preparation on the nonnull support. Values null for every preparation can be extended arbitrarily within their nonempty projection fibres. For padding invariance, reconstruct and append for one bound, and project away for the other.
□Independent noise and a copy of an already retained record satisfy this informational criterion. A coordinate whose marginal merely ignores the source need not satisfy it. Nor does snapshot redundancy imply that the coordinate is causally idle later: an otherwise redundant backup may become indispensable after the original register is overwritten.
4.3 The exact masked-encoder loss
The two rows in Equation 3.3, Equation 3.4 remain valid statistical experiments on the extended interval . Only the original graph argument was restricted to .
For the extended experiment,
The first formula changes branch at . Both losses are continuous and have explicit optimal reconstruction kernels.
For , the joint contrast is and the first-coordinate contrast is , so Equation 4.2 gives . For a second bound on the whole interval, put , . The reconstruction rows obey
Therefore . If both reconstruction errors are at most , the target masses give and . Hence . The two kernels in Appendix A attain the larger lower bound on the respective intervals. For output 2, the retained law is the same fair bit under both sources. Every decoder gives one common target law; the midpoint of the two targets attains half their TV diameter. Direct subtraction gives diameter .
□Figure 1 contrasts this continuous information loss with the inherited support selection. The new diagnostic does not reproduce the old graph while somehow removing its discontinuity: it deliberately asks a different, quantitative question.
4.4 Why a single TV increment is insufficient
No rule can discard every marginally source-independent coordinate while retaining all joint-only information. For a fixed contrast, or a common family of maximized contrasts,
but does not imply that some proper view reconstructs exactly.
Let with fair. Both marginals are independent of , while . Thus marginal independence cannot identify irrelevant padding. Nonnegativity of follows from contraction under projection, including after taking suprema over one common contrast family. At , the masked experiment has full and best-singleton contrasts both , yet Equation 4.9 gives and Equation 4.10 gives .
□There is a concrete decision witness. With source prior , optimal binary classification has Bayes error from the pair and from its first coordinate. The calculation is in Appendix A. Equal-prior pairwise discrimination is only one decision problem, so preserving its best score does not preserve the full experiment.
Redundant representations can also overlap. In , the pairs and each have deletion diagnostic , while . The full triple has not lost the bit: it contains redundant ways to retain it. A disjoint partition of informative representations is not implied by statistical sufficiency.