Skip to content
Shadow Theory

Paper 2 · Section 4Boundary

Recover the full experiment from a partial view

Directed statistical deficiency measures the information lost when an observation is deleted. Exact formulas, optimal decoders and a decision witness resolve the masked family across its full parameter interval.

Section 5 of 20

4 Robust reconstruction of joint experiments

4.1 The information lost by deleting a view

The obstruction concerns exact target selection, not the existence of stable measures of joint information. A different question is available: after retaining only part of an observation, can one common reconstruction procedure reproduce the full experiment for every unknown preparation?

Let Θ\Theta be a nonempty finite preparation set. An experiment has laws PJ,θP_{J,\theta} on a finite joint alphabet YJ\mathcal Y_J. A retained view is the deterministic projection πL:YJ→YL\pi_L:\mathcal Y_J\to\mathcal Y_L, with image alphabet YL\mathcal Y_L and marginal laws PL,θP_{L,\theta}.

Definition 4.1 (Directed reconstruction deficiency)

The loss of reconstructing JJ from LL is

δP(J∣L)=min⁡G:YL⇝YJmax⁡θ∈ΘTV⁡(PJ,θ,PL,θG). \delta_P(J\mid L)=\min_{G:\mathcal Y_L\rightsquigarrow\mathcal Y_J} \max_{\theta\in\Theta}\TV(P_{J,\theta},P_{L,\theta}G). (4.1)

The stochastic row kernel GG is the same for all unknown θ\theta. It may use a declared public context, but not the unknown source label.

This is directed statistical deficiency in a target-first, half-L1L^1 convention, not a new information measure. Statistical comparison has its classical Blackwell/Le Cam lineage [6, 16]; deficiency has also been applied directly to feature quality [24]. Argument order, row versus column kernels, and the factor of two in variational divergence must be translated when comparing formulas.

The decoder reproduces laws. In particular it may resample even the retained coordinates. It need not recover the same realized event, preserve a particular archive carrier, respect a deadline, or be available as a native physical operation. Those stronger requirements belong to the execution contract in Section 5, Section 6. They are not consequences of Equation 4.1.

Proposition 4.2 (Finite comparison bounds)

The minimum in Equation 4.1 exists, lies in [0,1][0,1], and is a rational linear-program value for rational finite data. For any θ,θ′\theta,\theta',

δP(J∣L)≥12[TV⁡(PJ,θ,PJ,θ′)−TV⁡(PL,θ,PL,θ′)]+. \delta_P(J\mid L)\ge \tfrac12\bigl[ \TV(P_{J,\theta},P_{J,\theta'})- \TV(P_{L,\theta},P_{L,\theta'})\bigr]_+. (4.2)

For experiments E,F,H\mathsf E,\mathsf F,\mathsf H on the same parameter set,

δ(E∣H)≤δ(E∣F)+δ(F∣H). \delta(\mathsf E\mid\mathsf H) \le \delta(\mathsf E\mid\mathsf F)+\delta(\mathsf F\mid\mathsf H). (4.3)

More informative retained observations cannot increase reconstruction loss, nor can postprocessing the target. If max⁡θTV⁡(PJ,θ,P^J,θ)≤ϵ\max_\theta\TV(P_{J,\theta},\widehat P_{J,\theta})\le\epsilon, with the same projection and preparation labels, then

∣δP(J∣L)−δP^(J∣L)∣≤2ϵ. \bigl|\delta_P(J\mid L)-\delta_{\widehat P}(J\mid L)\bigr|\le2\epsilon. (4.4)

All these statements preserve the nominated experiment and decoder permissions.

Proof

Finite stochastic matrices form a compact polytope and the objective is continuous. The absolute-value epigraph gives a finite linear program with rational coefficients; a rational optimum can be chosen. Insert PL,θGP_{L,\theta}G and PL,θ′GP_{L,\theta'}G between the two target laws and use TV contraction under GG to obtain Equation 4.2. Composing two reconstruction kernels proves Equation 4.3; the same argument gives the monotonicities. For a fixed GG, replace both target and input laws by their hatted versions. The first replacement costs at most ϵ\epsilon, and the second at most ϵ\epsilon after contraction. Take infima in both directions to obtain Equation 4.4.

□

For a decision loss valued in [0,1][0,1], any decision procedure based on JJ can be emulated from LL, using an optimal GG followed by that procedure, with additional risk at most δP(J∣L)\delta_P(J\mid L) for every preparation and hence every prior. This forward implication follows from the TV bound on bounded expectations. No full statistical-comparison converse is needed below.

One may record the entire table δP(J∣L)\delta_P(J\mid L). A derived deletion diagnostic is

uP(J)=min⁡j∈JδP(J∣J∖{j}), u_P(J)=\min_{j\in J}\delta_P(J\mid J\setminus\{j\}), (4.5)

where the empty view is a one-point alphabet. Positive uP(J)u_P(J) means that each one-coordinate deletion loses experiment information. It is not an additive information decomposition, a physical edge attribution, or an exhaustive notion of integration. In particular, it should not be identified with a partial-information-decomposition atom; unique-information constructions answer their own specified comparison questions [5].

4.2 Genuine informational padding

Proposition 4.3 (Zero loss and common-conditional padding)

For a deterministic view YL=πL(YJ)Y_L=\pi_L(Y_J), δP(J∣L)=0\delta_P(J\mid L)=0 if and only if there is one conditional kernel A(dyJ∣yL)A(dy_J\mid y_L), supported on πL(yJ)=yL\pi_L(y_J)=y_L, such that

PJ,θ=PL,θAfor every θ. P_{J,\theta}=P_{L,\theta}A\qquad\text{for every }\theta. (4.6)

In particular, if

Pθ(y,z)=Pθ(y)A(z∣y) P_\theta(y,z)=P_\theta(y)A(z\mid y) (4.7)

with the same conditional kernel in every preparation, then

δ((Y,Z)∣Y)=0,δ((Y,Z)∣L)=δ(Y∣L) \delta((Y,Z)\mid Y)=0, \qquad \delta((Y,Z)\mid L)=\delta(Y\mid L) (4.8)

for any common retained experiment LL.

Proof

A conditional reconstruction immediately proves the forward zero-loss claim. Conversely, an optimal zero-loss kernel reproduces the full experiment, though it need not initially preserve the view. Put a strictly positive prior on the finite Θ\Theta. The reconstructed Markov chain Θ→YL→YJ∗\Theta\to Y_L\to Y_J^* has the original (Θ,YJ)(\Theta,Y_J) law at its endpoints. Data processing gives I(Θ;YJ)≤I(Θ;YL)I(\Theta;Y_J)\le I(\Theta;Y_L). Since YLY_L is originally a deterministic function of YJY_J, the reverse inequality holds. Thus I(Θ;YJ∣YL)=0I(\Theta;Y_J\mid Y_L)=0 in the original experiment. Its conditional law supplies AA, common to every preparation on the nonnull support. Values null for every preparation can be extended arbitrarily within their nonempty projection fibres. For padding invariance, reconstruct YY and append AA for one bound, and project away ZZ for the other.

□

Independent noise and a copy of an already retained record satisfy this informational criterion. A coordinate whose marginal merely ignores the source need not satisfy it. Nor does snapshot redundancy imply that the coordinate is causally idle later: an otherwise redundant backup may become indispensable after the original register is overwritten.

4.3 The exact masked-encoder loss

The two rows in Equation 3.3, Equation 3.4 remain valid statistical experiments on the extended interval 0≤ε≤10\le\eps\le1. Only the original graph argument was restricted to 0≤ε≤120\le\eps\le\tfrac12.

Theorem 4.4 (Exact reconstruction loss of the masked encoder)

For the extended experiment,

δP(12∣1)=max⁡{12−ε,1−ε4},δP(12∣2)=12max⁡{ε,1−ε}.\begin{align}\delta_P(12\mid1)&=\max\left\{\tfrac12-\eps,\frac{1-\eps}{4}\right\},\tag{4.9}\\ \delta_P(12\mid2)&=\tfrac12\max\{\eps,1-\eps\}. \tag{4.10}\end{align}

The first formula changes branch at ε=13\eps=\tfrac13. Both losses are continuous and have explicit optimal reconstruction kernels.

Proof

For ε≤12\eps\le\tfrac12, the joint contrast is 1−ε1-\eps and the first-coordinate contrast is ε\eps, so Equation 4.2 gives 12−ε\tfrac12-\eps. For a second bound on the whole interval, put q+=(1+ε)/2q_+=(1+\eps)/2, q−=(1−ε)/2q_-=(1-\eps)/2. The reconstruction rows obey

Q0=q+G0+q−G1,Q1=q−G0+q+G1. Q_0=q_+G_0+q_-G_1,\qquad Q_1=q_-G_0+q_+G_1.

Therefore Q1(00)≥(q−/q+)Q0(00)Q_1(00)\ge(q_-/q_+)Q_0(00). If both reconstruction errors are at most tt, the target masses give Q0(00)≥12−tQ_0(00)\ge\tfrac12-t and Q1(00)≤tQ_1(00)\le t. Hence t≥(1−ε)/4t\ge(1-\eps)/4. The two kernels in Appendix A attain the larger lower bound on the respective intervals. For output 2, the retained law is the same fair bit under both sources. Every decoder gives one common target law; the midpoint of the two targets attains half their TV diameter. Direct subtraction gives diameter max⁡{ε,1−ε}\max\{\eps,1-\eps\}.

□

Figure 1 contrasts this continuous information loss with the inherited support selection. The new diagnostic does not reproduce the old graph while somehow removing its discontinuity: it deliberately asks a different, quantitative question.

Exact reconstruction-loss curves for the masking parameter from zero to one. The first loss is one half minus epsilon through epsilon one third, then one quarter of one minus epsilon. The second is one half of one minus epsilon through epsilon one half, then epsilon over two. Below, the inherited projection has both cross-edges at zero, but only the edge from 2 to 1 for positive epsilon through one half.
Figure 1. Minimal-target masking and quantitative deletion loss. The projected edge 1→21\to2 is present only at ε=0\eps=0 in the inherited family, whereas its joint reconstruction loss remains close to 1/21/2 near zero. The graph statement uses [0,1/2][0,1/2]; the displayed deficiency formulas are valid on [0,1][0,1]. The curves are exact formulas, not fitted data.

4.4 Why a single TV increment is insufficient

Proposition 4.5 (Marginal independence and scalar increments)

No rule can discard every marginally source-independent coordinate while retaining all joint-only information. For a fixed contrast, or a common family of maximized contrasts,

σ(J)=d(J)−max⁡L⊊Jd(L)≥0, \sigma(J)=d(J)-\max_{L\subsetneq J}d(L)\ge0,

but σ(J)=0\sigma(J)=0 does not imply that some proper view reconstructs JJ exactly.

Proof

Let (Y,Z)=(ξ⊕b,ξ)(Y,Z)=(\xi\oplus b,\xi) with ξ\xi fair. Both marginals are independent of bb, while Y⊕Z=bY\oplus Z=b. Thus marginal independence cannot identify irrelevant padding. Nonnegativity of σ\sigma follows from contraction under projection, including after taking suprema over one common contrast family. At ε=1/2\eps=1/2, the masked experiment has full and best-singleton contrasts both 1/21/2, yet Equation 4.9 gives δP(12∣1)=1/8\delta_P(12\mid1)=1/8 and Equation 4.10 gives δP(12∣2)=1/4\delta_P(12\mid2)=1/4.

□

There is a concrete decision witness. With source prior P(b=0)=3/4\Prb(b=0)=3/4, optimal binary classification has Bayes error 1/81/8 from the pair and 1/41/4 from its first coordinate. The calculation is in Appendix A. Equal-prior pairwise discrimination is only one decision problem, so preserving its best score does not preserve the full experiment.

Redundant representations can also overlap. In (Y1,Y2,Y3)=(ξ,ξ⊕b,ξ)(Y_1,Y_2,Y_3)=(\xi,\xi\oplus b,\xi), the pairs 1212 and 2323 each have deletion diagnostic 1/21/2, while uP(13)=uP(123)=0u_P(13)=u_P(123)=0. The full triple has not lost the bit: it contains redundant ways to retain it. A disjoint partition of informative representations is not implied by statistical sufficiency.