# Agency and Free Will

Jeremy Rodgers · EverythingEquation.com

Full published-paper treatment with separately identified supplementary inquiry.

Published source: https://doi.org/10.5281/zenodo.23202999


<!-- published: Section 1; /consciousness/agency/agency-as-a-physical-capacity -->

# Agency as a physical capacity

**A decision can be physically caused and still be governed by the agent's own evaluation.** The useful question is what the system actually does: which commitments it can inspect, which alternatives it can assess, which rule it can install, and how that rule governs the next decision.

Bounded Agency and Reflective Freedom turns that question into a set of finite constructions and exact tests. A controller reads retained commitments, assesses a represented decision rule, installs its assessment, and applies the rule to fresh inputs. Separate tests establish that its assessment agrees with its declared commitments and that the result actually travels through the claimed read and installation pathways. Information and resource bounds then determine how much of this capacity the controller can exercise.

This is the complete website treatment of Jeremy Rodgers's published paper, [Bounded Agency and Reflective Freedom: Evaluative Revision, Information Limits, and Faithful Realization](https://doi.org/10.5281/zenodo.23202999). The scientific account establishes **proximal evaluative self-governance**: present, locally implemented capacities for assessment and rule revision. Further pages develop the philosophical and technical implications without changing those results' assumptions.

## What an internal decision must demonstrate

An externally supplied answer, a stored command and an internal assessment can produce the same output. To distinguish them, we need more than a successful answer. We need an operating contract and tests of the mechanism.

The contract specifies information, time, executable operations and available resources. An alteration that an experimenter can physically impose is not automatically an option that the installed controller can execute. A controller might contain a dormant procedure yet have no way to call it before its deadline. It might represent an option that its hardware cannot carry out. The account measures what the actual architecture can recruit and do under the stated conditions.

Three obligations organize the work:

1. **Endorsement:** the assessment agrees with a relation specified in advance between retained commitments and acceptable amendments.
2. **Mediation:** actual internal reads and installation operations cause the tested result through the identified routes.
3. **Continuing disposition:** the installed rule governs later fresh cases, beyond the register value produced during revision itself.

These obligations are independent. A selector can be strongly sensitive to a retained bit and systematically choose the amendment that bit rejects. A controller can install an acceptable rule by accident while bypassing the retained commitment entirely. A one-time command overwrite can succeed without creating a new conditional disposition. Each fails a different part of the test.

## Three contributions, three distinct constructions

The first construction is an **editable charter**. Its finite grammar contains four conditional rules and two possible read orders. The controller assesses these charters against retained commitments, retains an already optimal rule, changes an inferior one, and uses the installed charter on a fresh case. A sharp information, mediation and endorsement inequality bounds what its inquiry transcript can support. A deliberately inverted selector shows why causal contrast alone cannot establish faithful evaluation.

The second contribution concerns **bounded inquiry**. Preservation criteria determine when the controller can continue making endorsed amendments while retaining its future audit capacity. Noisy-read bounds give exact limits on warranted revision. A separate parameterized sensor experiment solves the competition between inspecting three retained endorsements and calibrating an unknown reliability regime. Its theorem establishes the exact optimal adaptive performance and the first shared budget at which it exceeds every fixed-count allocation. The proof covers every admitted query interleaving, stopping rule and choice of unread endorsement index.

The third contribution is **faithful native realization**. A compiler preserves ordinary computations and their diagnostic interventions while adding protected-source returns at stated gate boundaries. Resource counts include the return schedule. A successor-revealing continuation establishes the required joint predictive congruence for the specified finite stochastic instruments. An access lemma ensures that native records do not give the executive an uncharged information channel. The same return construction also works for a constant-output controller, demonstrating that recurrence alone cannot supply evaluative competence.

These are related constructions with different guarantees. The compact charter evaluator revises a represented rule. The optimal inquiry machine allocates queries under a fixed majority charter. The noisy-amendment witness establishes a separate error frontier. Their contracts support comparison, but no single device is claimed to attain every bound simultaneously.

## Physical causation and reflective freedom

The philosophical position is compatibilist. A complete deterministic state can fix its actual continuation while the resulting mechanism still responds to relevant variations, assesses commitments and changes an operative rule. Reasons-responsiveness accounts already distinguish a mechanism's capacities from the history through which an agent comes to own that mechanism [6](/consciousness/agency/references#ref-6). The mathematical target here is the present organization, with bounded operations and an explicit retained-commitment relation.

The account allows model-free selection in a restricted domain. It also allows fixed metarules. A system can assess and change a represented rule without choosing the physical laws or every standard by which the assessment proceeds. Requiring prior self-authorization of every evaluative standard would add ultimate self-creation to the question. The demonstrated capacity is more precise: the present mechanism can inspect, retain and revise particular rules that govern later decisions.

Historical independence of values remains a further property. Neither internal location nor current responsiveness determines how a commitment was acquired. A successful assessment can still rely on false evidence or morally objectionable commitments. The tests distinguish fidelity to a declared evaluative relation from the truth of incoming evidence and the moral standing of the relation itself.

## How consciousness enters

In Shadow Theory, source and readout are operational aspects of a physical description. An aperture exposes only part of the source structure. This motivates precise accounting of available information, but incomplete access by itself establishes neither evaluation nor agency.

The SPC-2 consciousness constitution assigns experience to qualifying native recurrent organization under declared correspondence laws [19](/consciousness/agency/references#ref-19). Its treatments of boundaries, effective interfaces and interventional realization distinguish physical qualification from software behavior [18](/consciousness/agency/references#ref-18), [20](/consciousness/agency/references#ref-20), [21](/consciousness/agency/references#ref-21). This work asks whether an independently verified evaluator can exist in a common realization satisfying those native conditions. A decision score is never used as evidence sufficient for experience.

The resulting coexistence claim is conditional on the declared SPC-2 bridge laws. The evaluative mathematics can be assessed independently of those laws. Accepting the laws also does not make every recurrent controller evaluative: the constant-output counterexample preserves the distinction.

## Where the contribution sits

The paper makes its contribution through explicit constructions, sharp bounds and compatible realizations. It builds on established methods rather than treating each mathematical tool as a new principle.

| Earlier work | What it contributes | The specific task here |
| --- | --- | --- |
| List's analysis of AI free will [14](/consciousness/agency/references#ref-14) and Kenton and colleagues' causal identification of agents [13](/consciousness/agency/references#ref-13) | Intentional agency, alternative possibilities, causal control and adaptation under mechanism interventions | Verify that a declared bounded evaluator reads retained commitments, installs an assessed rule and uses it on fresh cases |
| Imperfect-information games [3](/consciousness/agency/references#ref-3) and certified self-modifying code [2](/consciousness/agency/references#ref-2) | Knowledge-set methods, preservation kernels and mutable instructions subject to invariants | Specify a relational evaluative property, actual amendment mechanism and information obligations needed to preserve it |
| Gödel machines and utility-sensitive self-modification [4](/consciousness/agency/references#ref-4), [24](/consciousness/agency/references#ref-24) | Assessment of modifications through retained standards | Test the actual bounded assessment and installation routes without requiring revision of every metarule |
| Metalevel decision processes [11](/consciousness/agency/references#ref-11) and restricted single-coin reliability models [15](/consciousness/agency/references#ref-15) | Selection of computations and the sensor law | Prove an exact calibration threshold and a shared-budget optimum with a paid-port implementation |
| Deadline-sensitive sequential testing [7](/consciousness/agency/references#ref-7) | Likelihood-ratio reasoning for fallible inquiry | Derive the exact finite endorsement-error frontier under the specified read law |
| Causal abstraction [1](/consciousness/agency/references#ref-1), [22](/consciousness/agency/references#ref-22) | Preservation of intervention laws under representation maps | Construct native schedules, records, physical refinements and access restrictions that meet the compatibility requirements |

A comparison class needs equal care. Fixed-order Boolean evaluation can stop when it finds a certificate even within a class called nonadaptive [12](/consciousness/agency/references#ref-12). The inquiry separation here concerns **fixed counts allocated across two tasks**. Its improvement comes from reallocating effort saved on one task to the other. That is the precise restriction the theorem removes.

This treatment begins with the operating contract, constructs rule revision, establishes its information and preservation limits, then develops the exact inquiry separation and its native realization. The philosophical conclusion follows the same route: freedom is expressed through available evaluative capacities and their limits in an embodied system.


<!-- published: Section 2; /consciousness/agency/operational-contract -->

# What self-governance requires

**A capacity belongs to an agent only when its installed organization can exercise it.** A represented possibility, an experimenter's intervention and a callable operation are different things. The operating contract makes those differences explicit before any claim of self-governance is tested.

## 2.1 State, information and executive targets

Fix a physical realization and a finite operational envelope. The complete state contains every persistent record, working register, installed program, phase, dispatch state and remaining resource that can affect the permitted future experiment. A native operation $u$ has the joint record and successor kernel

$$
K_u(x;r,x'),\qquad \sum_{r,x'}K_u(x;r,x')=1.
\tag{2.1}
$$

Boundary inputs and persistent source context must be included in $x$ or explicitly conditioned upon. Leaving out information that affects later transitions would change the model being certified.

The controller sees only the observations supplied by its installed interface. Its **executive information** can be smaller than the catalogue of native physical records used to certify a realization. A diagnostic intervention is an additional physical operation at a named occasion. It is not a free observation available to the ordinary decision procedure.

The deterministic special case is

$$
x_{t+1}=F(x_t,o_t),\qquad c_t=\pi(x_t,o_t).
$$

Evaluation, competition and installation are physical parts of this update. There is no additional chooser standing outside $F$. A selected command may operate an external actuator or change an internal target such as an attentional gate, retained priority or installed rule. An actuator failure can prevent the outward effect after selection has already occurred. Assessment and installation therefore need their own tests, separate from downstream control efficacy.

The contract names admissible preparations, interventions, native actions, resources and stopping conditions. Resource units are concrete: a one-bit query, an elementary gate, a joint acquisition interval and an individual physical write have different costs. Installing another observation port or permitting an uncharged parallel read changes the theorem's comparison class.

## 2.2 Executable options and actual decisions

A repertoire contains policies the present architecture can invoke or synthesize within a specified resource envelope. This requires more than physical possibility. It does not require every repertoire member to occur from the same fixed complete deterministic state.

An experimenter might overwrite a register to test its causal role even when the controller cannot perform that overwrite before its deadline. The intervention demonstrates sensitivity to the overwritten value. Endogenous availability requires an installed route by which the controller can reach it.

In the finite constructions, the admitted grammar consists of explicit programs or query trees with charged costs. **Active candidates** are represented in the present assessment. **Accessible candidates** can be recruited by an installed procedure before the remaining budget expires. A larger budget may support a different capacity. Duplicating command labels does not add options when the resulting consequences are equivalent.

### Definition 2.1: evaluative mediation

An assessment is evaluatively mediated under a declared contract when an installed procedure uses retained evaluative information to assess admitted alternatives, and a specified downstream disposition depends on that assessment through identified physical read and installation pathways. Agreement with an endorsement relation is checked separately from the causal effect of those pathways.

The endorsement relation states which amendments are acceptable given the retained commitments. It need not be a universal scalar utility. A scalar or lexicographic criterion suffices for the finite examples, while its physical presence establishes no moral correctness. The evaluative interpretation is supported by systematic use across candidate comparisons and fresh cases. Arbitrary influence from an internal bit would not establish that role.

### Definition 2.2: bounded represented-rule self-revision

A controller has bounded represented-rule self-revision when it can, within its declared budget, assess whether to retain or change a stored rule governing later decisions, install the result, and execute that rule on fresh admitted inputs. Retention after assessment is a legitimate result. The complete physical interpreter and some metarules may remain fixed.

An operative rule has become an object of assessment. This is a minimal reflective capacity. It does not require a model of every physical transition, autobiographical knowledge, general philosophical reflection or the ability to revise every foundational commitment. Each would be an additional capacity to establish.

## 2.3 The evaluative role

Baseline operational agency is an installed capacity to select among consequence-distinct executable transformations through retained evaluative organization. Learned or innate action tendencies can implement it. Simulating alternative futures is one possible implementation. Reflection adds assessment of the operative rule itself.

Learning from outcomes, planning, changing priorities and modifying a search procedure extend the baseline in different ways. Possessing one does not establish all the others. This restricted-domain account also does not conflict with world-model necessity results: Richens and colleagues recover environment models from competence across a specified goal family in stationary controlled Markov environments [17](/consciousness/agency/references#ref-17). The finite evaluator here is not claimed to possess that general competence.

A retained commitment has a comparative role. Across a declared set of situations, it changes which executable outcome or rule is favored. In the editable charter, physical considerations are already represented accurately and both actions are safe. The retained commitments determine how mixed cases should be settled. Revision changes a conditional priority rather than correcting a factual prediction.

This interpretation is fixed before the experiment. Its ranking, actual use and installed consequences receive separate tests. Defining all three after seeing the result would explain nothing. A sensor update can change a prediction; an arbitrary latch can change an output; neither alone establishes the declared ranking.

A simple engineered controller can satisfy the minimal operational criterion. Describing it as a machine does not refute the certificate, just as describing an event as internal does not establish it. A persistent one-bit memory lacks the installed evaluative selection among executive targets. The finite witnesses isolate a component of self-governance without claiming the full reflective life of a person.

### Assessment has no need for an infinite sequence of selectors

An assessment can take a represented rule as its object while leaving an interpreter or metarule fixed. Another admitted operation can assess some of those further structures. Requiring an autonomous prior choice of every criterion would change the target to ultimate self-creation.

The present claim is that an organized physical mechanism can assess, retain and change particular operative commitments. The preservation fixed points later in this treatment describe the continuing availability of that capacity. They do not introduce an evaluator outside physical causation.

### Inward acts, inhibition and tools

An executive target can be internal. Evaluation can redirect attention, inhibit an available action or install a revised rule. Mere storage of a sensory event is insufficient. The relevant selection or gating process must govern the update.

Inhibition need not be represented as a separate explicit command or assigned a universal additional energy cost. Its causal role is what needs establishing. These are possible executive targets; the finite witnesses specifically implement rule installation and command selection.

A tool's launch and its later operation are distinct causal stages. A one-way launch can express agency without telemetry. A jam can frustrate an attempt without erasing its prior assessment. Continuous feedback supports ongoing regulation but is unnecessary for every individual act. When a tool evaluates subordinate choices, proximal attribution may apply at both levels. Whether the tool belongs within an SPC-2 subject remains a separate physical boundary question.

## 2.4 Freedom has a profile

This account tracks three context-dependent, time-dependent capacities:

| Capacity | What must be established |
| --- | --- |
| Evaluative mediation | The agent's evaluative process governs the tested disposition |
| Executive options | Consequence-distinct operations can actually be recruited and executed |
| Assessment and revision | Operative rules can be examined, retained or changed within the available resources |

There is no asserted canonical scalar that combines these axes. Restraint can shrink an executive repertoire while leaving assessment intact. A previously adopted commitment can later obstruct revision. More output labels do not imply greater reflection.

Paired interventions, bounded decision trees, error probabilities and preservation sets measure parts of this profile in declared models. They do not rank every possible agent. Continued callability of an audit is required by the preservation theorems, not by every instance of agency. An intentionally irreversible act can still be agentic at the occasion of its selection.

Three distinctions keep the profile operational.

**Representation and execution need a realization map.** An overlooked exit may be physically available but unrecruited. An imagined escape may be impossible. Before comparing repertoires, connect represented options to policies the architecture can execute at the relevant state and budget. An impossible counterfactual may help thought without adding an executable option.

**Stored possibilities require a dispatcher.** A memory entry becomes accessible only when an installed procedure can retrieve, parameterize and assess it before the deadline. An analyst's ability to invoke a dormant subroutine is not the agent's ability to do so.

**Consequence diversity needs a horizon.** Specify the observed variables, time horizon, risks and future continuation opportunities. Otherwise redundant micro-command encodings can inflate the apparent number of choices.

Revision has the same temporal structure. An audit may be underway, callable before the present deadline, or feasible only under a larger future budget. A high-friction review procedure can act as a practical lock now even if it succeeds later. Charging inspection and installation gives this distinction exact content.

Structural reconfiguration can expand a search grammar, but expansion is not the only improvement. Removing a redundant branch may make deliberation more efficient. Precommitment can protect one valued trajectory and foreclose another. Changes must be assessed on the affected axes and occasions. An earlier exercise of self-governance does not automatically increase present freedom. The finite charter changes a stored rule and read order; the validator construction changes audit modes. Neither establishes unrestricted cognitive self-redesign.

For sentient agency, qualifying SPC-2 organization is an additional constitutive requirement. Evaluative control and experiential assignment are different predicates. Their coexistence needs a common realization, which the later native construction provides conditionally.

## 2.5 Testing the actual causal pathways

Take two admitted preparations with identical present external query, phase and resources, but different retained evaluative states. Let $P_0,P_1$ be the laws of a downstream record that reveals the assessment or installed disposition. Apply a common intervention on the actual read route and, separately, on the actual installation route.

The scored record begins **after** the intervention and excludes a direct copy of the pre-intervention archive. A surviving historical copy would otherwise look like a bypass of the tested route. A positive intact contrast that falls under an actual cut certifies the tested mediated contribution. It does not establish that all historical influence uses that route.

### Proposition 2.3: transport of a finite certificate

Suppose state and record bijections transport the preparation laws, joint kernels, diagnostic interventions, executive observation maps, resource charges and stopping rule of a finite experiment. Every intact and intervened transcript law is then transported exactly. Event scores and total-variation contrasts are invariant.

**Proof.** The transported joint kernels give the one-step equality. Conditional on corresponding histories, the transported executive chooses corresponding actions. Multiplying by the next joint kernel and summing over corresponding states preserves the joint history/state law. Diagnostic interventions satisfy the same step. Induction reaches the bounded stopping horizon. Bijections preserve event probabilities and total variation. $\square$

This compatibility principle agrees with causal abstraction [1](/consciousness/agency/references#ref-1), [22](/consciousness/agency/references#ref-22). An equivalent lookup implementation is not disqualified because it looks less deliberative. Ordinary input-output agreement is still insufficient: it can coexist with incorrect intervention behavior or extra information access. Mixing components or changing the physical subject boundary needs its own realization argument. Proposition 2.3 does not assert SPC-2 invariance under those changes.

### Finite implementation error

If corresponding joint kernels differ uniformly by at most $\epsilon_t$ in total variation, and the initial laws differ by at most $\epsilon_0$, a stepwise coupling gives

$$
\operatorname{TV}(P_{\mathrm{trace}},\widetilde P_{\mathrm{trace}})
\leq \eta:=\min\left\{1,\epsilon_0+\sum_t\epsilon_t\right\}.
\tag{2.2}
$$

Until the first mismatch, the adaptive actions coincide. A union bound on that first mismatch proves the inequality. With common per-arm error $\eta$, an intact-minus-cut contrast can fall by at most $4\eta$, and a two-arm expected-score advantage can fall by at most $2\eta$.

These are guarantees on a finite experiment's **joint record and successor law**. Output marginals alone do not control them. Faithful realization preserves the machinery that makes the certificate meaningful, including what the executive can and cannot observe.


<!-- published: Section 3; /consciousness/agency/revising-an-evaluative-rule -->

# A rule can become an object of assessment

**An agent can change the rule governing its later decisions without changing the laws governing its physical implementation.** The editable-charter construction demonstrates this capacity in full. Its archive supplies commitments, its auditor compares candidate rules, and its installation mechanism changes a conditional disposition that is then tested on fresh inputs.

Self-modifying code and utility-sensitive self-modification already have formal treatments [2](/consciousness/agency/references#ref-2), [4](/consciousness/agency/references#ref-4), [24](/consciousness/agency/references#ref-24). The construction here makes the evaluative role and its actual causal pathways independently testable. Selecting an action and revising the rule that will select later actions become distinct operations with distinct records.

## 3.1 The charter

Let $e=(e_0,e_1)\in\{0,1\}^2$ represent two considerations whose relevant physical consequences are accurately represented. Both final actions are executable and safe. A writable charter is

$$
\chi=(q_{01},q_{10},d)\in\{0,1\}^3.
$$

Its conditional rule is

$$
\begin{aligned}
f_\chi(0,0)&=0,& f_\chi(1,1)&=1,\\
f_\chi(0,1)&=q_{01},& f_\chi(1,0)&=q_{10}.
\end{aligned}
\tag{3.1}
$$

The mixed-case outputs specify four possible rules:

| $q_{01}$ | $q_{10}$ | Rule |
| --- | --- | --- |
| $0$ | $0$ | AND |
| $1$ | $1$ | OR |
| $0$ | $1$ | Projection onto $e_0$ |
| $1$ | $0$ | Projection onto $e_1$ |

The bit $d$ sets which consideration is read first. Execution reads $e_d$ and exits immediately if both completions of the unread consideration give the same output. Otherwise it reads the other bit. The consideration-read cost is

$$
r(\chi,e)\in\{1,2\}.
$$

The operative priority register $\varrho$ receives an answer only when ordinary execution later occurs. Revising the charter is a separate event.

## The retained commitments

The retained record is

$$
z=(a,b,s,h_0,h_1).
$$

Here $a$ and $b$ endorse outputs for the two mixed-consideration cases. The bit $s$ records a commitment to symmetry between those cases. The pair $h=(h_0,h_1)$ stores a previous ordinary case.

For $q=(q_{01},q_{10})$, the installed metarule uses the loss

$$
L_z(q)=\mathbf 1\{q_{01}\ne a\}
+\mathbf 1\{q_{10}\ne b\}
+2s\,\mathbf 1\{q_{01}\ne q_{10}\}.
\tag{3.2}
$$

The symmetry weight and the interpretation of the endorsements are explicit premises. Physical dynamics do not derive their moral authority. What matters for this experiment is that the controller evaluates the stored commitments instead of receiving an external answer that names the new program.

Given the current charter $\chi^{\mathrm{old}}=(q^{\mathrm{old}},d^{\mathrm{old}})$, the auditor scans all eight charters and minimizes the following key lexicographically, comparing its first component first and moving right only to resolve ties:

$$
\begin{aligned}
\bigl(&L_z(q),\ \mathbf 1\{q\ne q^{\mathrm{old}}\},\ r(\chi,h),\\
&\mathbf 1\{d\ne d^{\mathrm{old}}\},\ q_{01},q_{10},d\bigr).
\end{aligned}
\tag{3.3}
$$

The priorities have clear operational meaning. First minimize disagreement with the retained commitments. Retain the current conditional rule if it is equally endorsed. Then minimize reads on the retained case. Preserve the old read order if it is equally economical. The remaining bits provide a definite tie-break.

The read-order improvement is empirical and local. One stored case gives no generalization guarantee for an unknown future distribution.

### Proposition 3.1: selective represented-rule revision

For every initial charter and retained record, the scan:

- installs a minimum-loss charter;
- retains the current conditional rule whenever that rule has minimum loss;
- minimizes retained-case read cost subject to the preceding requirements;
- leaves $\varrho$ unchanged throughout revision;
- returns to the same callable audit interface;
- preserves exact ordinary execution using at most two consideration reads.

Any changed conditional rule differs from its predecessor on a fresh mixed-consideration case.

**Proof.** After every scan step, the saved candidate has the smallest key among the candidates examined so far. All eight candidates are examined, which proves optimization and retention. Installation writes only charter registers. Resetting scratch while retaining the entry point leaves the operative priority untouched and makes the audit callable again. Ordinary execution exits after one read only when both completions agree; otherwise its second read determines the full input. Finally, different conditional rules differ in $q_{01}$ or $q_{10}$, so one of the two fresh mixed cases separates them. $\square$

![Functional organization of the revision test: retained commitments feed assessment, assessment installs a rule, and the installed rule acts on fresh considerations to produce the tested disposition.](/publications/consciousness/agency/figures/revision-test.svg)

*Figure 1. Retained commitments feed assessment; assessment installs a rule; that rule acts on fresh considerations. A read intervention acts on the commitment-to-assessment route. An installation intervention acts on the rule write. These are functional roles, not the primitive physical component graph used for SPC-2 qualification.*

## A changed disposition, demonstrated on a new case

Start from AND with $d=0$ and $\varrho=0$. The retained record

$$
z=(1,1,1,0,1)
$$

selects OR and changes the first read to $d=1$. The priority $\varrho$ remains zero during this revision. A later input $(0,1)$ then produces one, where the former AND rule would have produced zero.

The change concerns a conditional disposition and its read order. There was no factual correction and neither action was unsafe. Fixed physical instructions carry out the assessment and installation. The controller's total transition law remains fixed while its represented rule changes.

### Every stage has a cost

Eight candidate assessments are not eight elementary gates. A direct serial implementation charges five record reads, eight key evaluations and comparisons, candidate storage, installation and rearming.

The logical reference implementation has exact minimum worst-case archive-read depth **five** when the initial charter is AND or OR, and **four** when it is a projection. These depths assume one-bit access and no advice about unread records. The explicit query recurrence and exhaustive finite check appear in Appendix A.1.

A circuit that reads the entire archive can preserve the functional result of Proposition 3.1 while losing its short-circuit timing. Matching a rule does not automatically preserve every resource guarantee of another implementation.

## 3.2 Testing reading, installation and endorsement

Compare two matched preparations: one has $(a,b)=(0,0)$ and the other $(a,b)=(1,1)$. Hold the symmetry bit, retained case, initial charter, current priority, resources and fresh probe fixed.

The first preparation installs AND; the second installs OR. The later probe $(0,1)$ separates them. A common clamp on the **actual endorsement read route** removes that contrast. Holding the **actual charter-write gate** fixed also removes the later probe contrast even if the internal candidate assessments still differ. Clamping an external actuator, by contrast, need not erase the installed internal disposition.

These interventions distinguish four events: reading, assessing, installing and successfully acting outward. They establish causal roles. A further check establishes whether the installed rule agrees with the retained endorsements.

### Proposition 3.2: information, installed contrast and endorsement

Suppose two matched retained contexts endorse disjoint amendment sets $A_0,A_1$. Let $H_i$ be the actual inquiry-transcript law in context $i$, and suppose the installed amendment law is

$$
P_i=H_iK,
$$

with the same downstream Markov kernel $K$ in both contexts. Set $\epsilon_i=P_i(A_i^c)$, the probability of an unendorsed installation. Then

$$
\max\{0,1-\epsilon_0-\epsilon_1\}
\leq\operatorname{TV}(P_0,P_1)
\leq\operatorname{TV}(H_0,H_1).
\tag{3.4}
$$

In particular,

$$
\epsilon_0+\epsilon_1\geq1-\operatorname{TV}(H_0,H_1).
$$

**Proof.** Because $A_0$ and $A_1$ are disjoint, $P_0(A_1)\leq\epsilon_0$ while $P_1(A_1)\geq1-\epsilon_1$. Testing the event $A_1$ gives the lower bound. Total variation contracts under the common kernel $K$, giving the upper bound. $\square$

The common-kernel condition excludes an uncharged route carrying the retained context around the inquiry interface. Installation noise is allowed inside $K$. A downstream external probe inherits the lower bound only if it distinguishes the relevant installed amendment classes, as the mixed-case probe does.

The inequality puts a precise limit on evaluation. If two contexts require disjoint amendments yet the permitted inquiry barely distinguishes them, low endorsement error is impossible. Internal processing cannot recover a distinction that never reaches it through the allowed route.

## A causally responsive evaluator can still be wrong

The bound is sharp. Consider a binary read with crossover probability $\nu<1/2$.

| Selector | Endorsement error | Intact contrast | Contrast under a common read or installation clamp |
| --- | --- | --- | --- |
| Install the reported charter | $\nu$ | $1-2\nu$ | $0$ |
| Install the opposite charter | $1-\nu$ | $1-2\nu$ | $0$ |

With noiseless reads, both selectors have contrast one. One always installs an endorsed charter; the other never does. Their causal sensitivity is identical.

**Mediation locates a causal route. Endorsement tests what that route accomplishes.** The certificate needs both. Conversely, one accidentally endorsed installation cannot establish that retained commitment mediated the decision. This distinction gives the finite account more content than an output score or a generic claim that internal state matters.


<!-- published: Section 4; /consciousness/agency/preserving-revision -->

# Keeping revision available

**A safe amendment may exist while the agent cannot identify and install it in time.** Preserving reflective capacity therefore requires a joint account of what remains endorsed, what preserves the audit mechanism and what the present controller can find through its actual interface.

Let $P$ be a finite set of admitted controller configuration types, and let $G\subseteq P$ contain the configurations satisfying a specified local audit certificate. The certificate includes bounded completion, agreement with the declared evaluative relation, and paired preparation and read/commit controls. It is evaluated on transition semantics. A register that merely says “certified” is insufficient.

Configuration types may include writable assessment procedures and validators. The interpreter semantics and the property being proved remain fixed. Revising a validator cannot establish preservation simply by changing what success means.

## The information available at an audit

At configuration $p$, a retained context $z\in Z_p$ is fixed throughout the current audit, with $Z_p$ finite and nonempty. The relation

$$
N(p,z)\subseteq P
$$

specifies endorsed amendments. Every configuration type and its context model must retain all history, legality and resource information needed for future transitions. Successive contexts must compose under the stated envelope.

A fresh budget at every audit is a replenishment assumption. If the controller instead has a finite lifetime reserve, that reserve belongs in its state. Otherwise an apparently repeatable guarantee could silently spend the same resource more than once.

After an internal transcript $h$, write $Z_p(h)$ for the contexts still compatible with the observed history. Installation has no additional access to the unread context.

### Lemma 4.1: a common amendment

For a proposed continuation set $X\subseteq G$, a definite zero-error installation into $X$ is possible exactly when the installed dispatcher has a budget-feasible installation kernel supported entirely on

$$
\bigcap_{z\in Z_p(h)}\bigl[N(p,z)\cap X\bigr].
\tag{4.1}
$$

For deterministic installation, the intersection must contain an amendment that the dispatcher can actually execute with its remaining budget.

**Proof.** Every amendment in the support of such a kernel is endorsed and belongs to $X$ in every compatible context, giving probability-one success. Conversely, any positive output mass outside the intersection causes error in at least one compatible context. Finiteness makes the support condition necessary for randomized installation too. The dispatcher must realize this supported kernel within budget; having a positive chance of reaching one acceptable amendment is insufficient. $\square$

The intersection captures the cost of unresolved information. When the same observed history remains compatible with contexts that demand incompatible amendments, the controller needs more information before it can complete without error. Mere existence of an acceptable amendment in each separate context does not solve the problem.

## 4.1 The preservation set

Let $C_B(p,X)$ mean that an admitted internal query-and-installation tree completes within budget $B$ and satisfies Lemma 4.1 at every leaf. Its code, working memory, branch operations and installation costs are all included in the machine contract.

Start with the locally certified configurations and repeatedly remove those unable to complete an endorsed audit back into the surviving set:

$$
\begin{aligned}
W_0&=G,\\
W_{j+1}&=\{p\in W_j:C_B(p,W_j)\},\\
W_\infty&=\bigcap_{j\geq0}W_j.
\end{aligned}
\tag{4.2}
$$

Every admitted tree must be callable by the current installed dispatcher. Computing the fixed point does not authorize an experimenter to replace that dispatcher. If the contract allows synthesizing and installing a new tree, those operations also need executable routes and charged costs. Composition includes the storage and dispatch needed to implement the selected trees together.

### Theorem 4.2: budgeted revision preservation

Assume:

- contexts remain stable during each audit;
- the operational model is finite, exact and composable;
- the charged policy class is closed under the stated continuation composition.

Then $W_\infty$ is exactly the set from which an admitted bounded revision procedure can honor endorsement and preserve $G$ at every future audit occasion. The iteration stabilizes after at most $|P|$ strict removal stages. An attaining tree at each surviving configuration realizes the guarantee.

**Proof.** The predicate $C_B(p,X)$ is monotone in $X$: a tree supported on an acceptable set remains acceptable if the continuation set expands. At the finite fixed point, every leaf of an attaining tree installs an endorsed successor in that same fixed point within budget. Induction over audit occasions proves sufficiency.

Conversely, let $Y$ be the configurations from which the stated universal guarantee is achievable. The first bounded audit at any $p\in Y$ must complete in $Y$ on every compatible branch. Otherwise a continuation could violate the guarantee. Lemma 4.1 therefore gives $C_B(p,Y)$. Since $Y\subseteq G$, induction gives $Y\subseteq W_j$ for every $j$, hence $Y\subseteq W_\infty$. Each strict stage removes at least one member of finite $P$, proving the stage bound. $\square$

This is a partial-observation safety construction using established fixed-point and decision-tree methods [3](/consciousness/agency/references#ref-3). Its contribution to the agency account is the explicit distinction between an acceptable amendment's existence and the installed controller's ability to discover and execute it.

## One tick changes the answer

Suppose a source configuration contains an unread bit $z$. Only a left amendment is endorsed when $z=0$, and only a right amendment is endorsed when $z=1$. Both successor configurations preserve the audit certificate. Reading costs one tick; installation costs another.

With a one-tick budget, the source is excluded from $W_\infty$. Its interface cannot both distinguish the contexts and install the appropriate successor. With two ticks, it is admitted. The set of acceptable successors under full information is identical in both cases. What changes is executable access before the deadline.

## Preservation is a chosen objective

The theorem tracks continued auditability. It does not require every agentic act to preserve all future options. An endorsed irreversible commitment can conflict with the preservation objective while remaining an exercise of agency at its own occasion.

Writable validators also fit the framework. Their transition semantics must preserve the original certificate. Merely checking that the next program passes its present local test can miss a later loss of responsiveness. Appendix A gives the explicit query-cost recurrence, a universal closure kernel and a finite example with mutable validators that separates local certification from continued preservation.


<!-- published: Section 5; /consciousness/agency/fallible-internal-inquiry -->

# When inquiry is fallible

**More inquiry can reduce the probability of a wrong amendment without making error impossible.** A bounded audit needs to distinguish these claims. It also needs to distinguish choosing an unendorsed rule from losing the capacity to audit again.

Deadline-sensitive inference and risk-constrained partially observed planning supply established tools for this problem [7](/consciousness/agency/references#ref-7), [23](/consciousness/agency/references#ref-23). Here the loss is disagreement with an internally retained endorsement under a declared metarule. The benchmark makes that loss exact.

## 5.1 Why finite noisy inquiry cannot guarantee certainty

### Proposition 5.1: the finite full-support obstruction

Assume finite response alphabets and a finite set of initially compatible contexts. Every admitted query response has common positive conditional support across those contexts. Query choice, timing and availability depend only on observable history, and all context information reaches installation through that history. Every audit terminates within a finite worst-case operation bound.

A definite zero-error endorsed amendment into $X$ then requires

$$
\bigcap_{z\in Z_p}\bigl[N(p,z)\cap X\bigr]\ne\varnothing.
\tag{5.1}
$$

A common member that can be installed with zero error within budget is sufficient. Finite inquiry cannot repair an empty intersection.

**Proof.** Every reachable terminal transcript has positive probability under every initially compatible context. Along that transcript, the policy chooses the same queries and each response retains positive conditional probability. The installed amendment must consequently be valid in all the original contexts. For randomized policies, apply the same reasoning to their common randomization kernel and terminal amendment support. Direct installation of a common member proves sufficiency. $\square$

The assumption concerns positive probability, not merely topological support. For more general response spaces, mutual absolute continuity of the conditional response laws along common histories is a corresponding sufficient hypothesis. If installation itself is noisy, its output support must also meet Lemma 4.1's zero-error condition.

The result does not say noisy inquiry is useless. Context probabilities can change dramatically while every context remains possible. To measure that change, we need a risk frontier.

## 5.2 A noisy charter audit

Restrict the charter to

$$
a=b=z\in\{0,1\},\qquad s=1,
$$

with a fixed retained case and read-order convention. The endorsed programs are $p_0$ (AND) and $p_1$ (OR), each preserving the same audit interface. The only admitted access to the retained bit is

$$
Y_i=z\oplus N_i,\qquad
N_i\overset{\mathrm{iid}}{\sim}\operatorname{Bernoulli}(\nu),
\qquad0<\nu<\frac12.
\tag{5.2}
$$

The independence law is conditional on the retained state and relevant past. There is no noiseless replica, initial advice, confidence flag or uncharged side channel.

### Theorem 5.2: the exact bounded noisy-revision frontier

Allow at most $n\geq1$ reads, adaptive stopping and declared randomization. Require definite installation of $p_0$ or $p_1$ before the deadline. The minimum worst-case endorsement error is

$$
\begin{aligned}
R_n(\nu)={}&\sum_{k>n/2}\binom nk\nu^k(1-\nu)^{n-k}\\
&+\frac12\mathbf1\{n\text{ even}\}\binom n{n/2}
[\nu(1-\nu)]^{n/2}.
\end{aligned}
\tag{5.3}
$$

A deterministic majority of the largest odd number of reads at most $n$ attains the bound. Loss of the specified future audit capacity has probability zero.

**Proof.** Give the two retained states equal prior probability. After $n$ reads, with $k$ observed ones, the likelihood ratio is

$$
\left(\frac{1-\nu}{\nu}\right)^{2k-n}.
$$

The minimum average error is majority's binomial error, with half the tie mass, yielding (5.3). A policy supplied all $n$ reads can simulate every earlier-stopping policy by ignoring an unused suffix. The Bayes error is therefore a lower bound on every admitted policy's worst-case error.

For odd $n$, majority has exactly that error in both retained states. Binomial expansion gives $R_{2k}=R_{2k-1}$, so an even cap is attained by majority on the preceding odd number of reads, without needing a random tie-break. Both possible installations preserve the audit interface by construction. $\square$

The result separates an error in the installed rule from a loss of reflective capacity. A fallible amendment can leave the next audit fully available. Neither score should stand in for the other.

## The cost of a warranted tolerance

The attaining policy uses a signed balance and a read counter:

$$
S\leftarrow S+2Y_i-1.
$$

The balance needs $\lceil\log_2(2n+1)\rceil$ bits and the counter needs $\lceil\log_2(n+1)\rceil$ bits, beyond latches, program and phase storage.

Let $c_{\mathrm{read}}$ charge a complete read-and-update step. Let $c_0$ charge initialization, comparison, installation and rearming. The tolerance cost is

$$
B_\epsilon=c_0+c_{\mathrm{read}}
\min\{n\geq1:R_n(\nu)\leq\epsilon\}.
\tag{5.4}
$$

These are macro-operation costs, not elementary gate counts. For $\nu=1/4$, the exact frontier gives:

| Endorsement-error tolerance | Required reads |
| --- | --- |
| $10^{-1}$ | $7$ |
| $10^{-2}$ | $19$ |
| $10^{-3}$ | $33$ |
| $10^{-4}$ | $49$ |

With no reads, a deterministic forced decision has worst-case error one. Error one half at zero reads requires an explicitly available fair random source. Randomization is an operation with a stated realization, not an implicit gift to the controller.

## What the intervention test measures

For matched preparations $z=0,1$, a fresh mixed-case probe has intact contrast

$$
1-2R_n.
$$

A common read clamp and a common commit clamp each reduce the contrast to zero. At $\nu=1/4$ and $n=3$:

$$
R_3=\frac5{32},\qquad
\text{intact contrast}=\frac{11}{16},\qquad
\text{capacity-loss probability}=0.
$$

The three numbers answer different questions: how often the amendment disagrees with retained endorsement, how strongly the retained commitment changes the later disposition, and whether another audit remains possible.

## Repeated audits and their joint risk

For $L$ fresh-context audits with the same read cap and the stated independent conditional read law, the exact minimax probability of at least one unendorsed installation is

$$
1-(1-R_n)^L.
\tag{5.5}
$$

Majority attains it by multiplying conditional success probabilities. For the converse, independently uniform fresh retained bits give every policy conditional success at most $1-R_n$ at each audit. The probability of succeeding throughout is therefore at most $(1-R_n)^L$.

A persistent hidden commitment would define a different model. Earlier inquiries could remain informative, so one could not reset the uncertainty at every audit. The general risk-vector treatment in Appendix A retains the information needed to compose the appropriate experiment.

## Two ways to misread an improvement

**Deferral changes the task.** If the controller may defer, report both deferral probability and error conditional on completion. A policy that never finishes makes no erroneous installation, but has not solved definite revision before a deadline.

**Repeated errors may be correlated.** If every read shares one noise bit,

$$
Y_i=z\oplus N,
$$

each marginal read still has error $\nu$, yet the entire transcript contains only one observation. The minimax error stays $\nu$. Repetition improves the frontier under conditional independence, not under marginal accuracy alone.

Bounded inquiry gives reflection a measurable limit. A controller can be genuinely responsive to its commitments, can preserve its capacity to reconsider, and can still lack the information needed to guarantee that the present amendment is endorsed. That combination is a feature of finite evaluation, not a contradiction in it.


<!-- published: Section 6; /consciousness/agency/evidence-and-adaptive-inquiry -->

# Evidence and adaptive inquiry

An agent has two jobs before acting: find out what is happening, and inspect what it has committed to doing. Those jobs compete for the same finite time. A useful discovery in one can change how much work remains for the other. The question is whether acting on that discovery gives a provable advantage over deciding the allocation in advance.

Here the answer is exact. In a specified sensor family, the first useful calibration count is $L$. The first shared budget at which adaptive allocation beats every fixed-count allocation is **$L+2$**. The stronger worked example needs four inquiry opportunities. Its advantage comes from using the effort saved by an early endorsement certificate to cross a factual decision threshold.

This chapter contains the complete model and proof from Section 6 of the published paper. The source law specializes a single-coin reliability model [15](/consciousness/agency/references#ref-15), and the allocation problem belongs to metalevel decision theory [11](/consciousness/agency/references#ref-11). Its reward combines factual accuracy and agreement with retained commitments. Neither component is a test of moral rightness.

## The source and the clock

A stationary hidden regime $\theta\in\{0,1\}$ has equal prior probabilities. The three sensor skills are

$$
a^{(0)}=(u,v,w),\qquad a^{(1)}=(v,u,w),\qquad
0<u<v<1,\quad 0<w<1,\quad v>\frac{u+w}{1+uw}.\tag{6.1}
$$

Each fresh item has an independent fair label $X\in\{-1,+1\}$. Its outputs are $Y_i=XE_i$. Conditional on the regime, errors are independent across sensors and items, with $\mathbb E[E_i\mid\theta]=a_i^{(\theta)}$. A skill $a$ therefore means accuracy $(1+a)/2$. The complete joint law is

$$
f_\theta(x,y):=\Pr_\theta(X=x,Y=y)
=\frac12\prod_{i=1}^3\frac{1+xy_i a_i^{(\theta)}}2.\tag{6.2}
$$

The controller knows this family and the positive orientation of the skills. It receives neither the regime nor the true labels of calibration items. Independence, stationarity, orientation and restriction to this family are assumptions. Unlabeled agreement cannot establish them all: reversing every skill and every hidden label preserves the observable law while reversing the correct prediction.

Three retained endorsement signs $e_1,e_2,e_3$ are independent and fair in the benchmark distribution, and independent of the source. The installed charter requires the operative priority to equal $\operatorname{maj}(e_1,e_2,e_3)$. Priority has its own register. Installing the majority can retain or revise that priority, but **the majority charter stays fixed in this experiment**. Changing the charter itself is the separate [editable-rule construction](/consciousness/agency/revising-an-evaluative-rule).

There are at most $B$ inquiry opportunities. One opportunity buys either a complete raw sensor triple on a fresh calibration item or one previously unread endorsement. The controller can choose the operation and named endorsement index from its entire acquired history and an independent random seed; it may stop early. The final target triple arrives only after inquiry. It cannot guide earlier requests. Success requires both a correct target prediction and a correct installed endorsement majority.

The initially accessible state, including the old priority, is independent of regime and endorsements. There is no initial advice, timing signal or extra output port that reveals them. Without that restriction the problem changes: an already visible correct majority would give zero-query joint success $b$, rather than $b/2$.

A query unit compares information access. It is not a universal energy unit. Computation, memory, ingress, routing and clocks receive separate charges in the [realization](/consciousness/agency/faithful-native-realization). A physically retained record does not automatically give the executive a free read.

## Information can improve before decisions do

Compress each calibration triple to a signed increment:

$$
Z(y)=\begin{cases}
+1,&y_1=y_3\ne y_2,\\
-1,&y_2=y_3\ne y_1,\\
0,&\text{otherwise},
\end{cases}
\qquad S_n=\sum_{j=1}^n Z(Y^{(j)}).\tag{6.3}
$$

Summing the source law over its hidden label gives

$$
\Pr_\theta(Y=y)=\frac{1+a_1a_2y_1y_2+a_1a_3y_1y_3+a_2a_3y_2y_3}{8},
$$

where the skills are those of the chosen regime. Put

$$
\begin{aligned}
p_-&=\frac{1-uv+w(v-u)}4,&p_+&=\frac{1-uv-w(v-u)}4,\\
p_0&=\frac{1+uv}2,&\rho&=\frac{p_-}{p_+},\\
A&=v-u-w(1-uv),&D&=v-u+w(1-uv).
\end{aligned}\tag{6.4}
$$

Under regime zero, $(Z=-1,0,+1)$ has probabilities $(p_-,p_0,p_+)$. Under regime one, the outer probabilities exchange places. Every raw triple with $Z=0$ has the same probability in both regimes. Thus every complete calibration history has regime-one to regime-zero likelihood ratio $\rho^{S_n}$, giving the exact posterior

$$
\pi(s):=\Pr(\theta=1\mid S_n=s)=\frac{\rho^s}{1+\rho^s}.\tag{6.5}
$$

No supplied confidence flag enters this calculation. Let $b$ be target-majority accuracy and $c$ the accuracy available if the regime were revealed:

$$
b=\frac{2+u+v+w-uvw}{4},\qquad c=\frac{1+v}{2}.\tag{6.6}
$$

The dominance condition makes the stronger sensor's log-odds weight exceed the other two combined. Following that sensor attains $c$.

### Proposition 6.1: the sharp calibration threshold

Before seeing the target triple, optimal expected target accuracy conditional on calibration score $s$ is

$$
M(s)=\max\left\{b,\frac{1+u}{2}+\frac{v-u}{2}\max\{\pi(s),1-\pi(s)\}\right\}.\tag{6.7}
$$

Define

$$
L=\min\{n\ge1:\rho^n>D/A\}.\tag{6.8}
$$

Then $L\ge2$, $b\le M(s)\le c$, and $M(s)=b$ for $|s|<L$. Exactly $L$ calibrations first permit strict improvement. Writing $C_L=\mathbb E[M(S_L)]$,

$$
C_L-b=\frac{Ap_-^L-Dp_+^L}{4}>0.\tag{6.9}
$$

An attaining classifier follows sensor one for $s\ge L$, sensor two for $s\le-L$, and target majority otherwise. At an exact likelihood threshold it may retain majority.

**Proof.** For each target cue $y$, its signed truth margin is

$$
d_\theta(y):=f_\theta(+1,y)-f_\theta(-1,y)
=\frac{\sum_i a_i^{(\theta)}y_i+uvw\,y_1y_2y_3}{8}.
$$

At posterior odds $\ell$, the Bayes action has the sign of $d_0(y)+\ell d_1(y)$. For $y=(+1,-1,+1)$ the margins are $-A/8$ and $D/8$. Swapping the first two sensors or reversing all signs generates the remaining disagreement cases. The only positive switching thresholds are $A/D$ and $D/A$. If the first two sensors agree, their sign is optimal in both regimes: the potentially competing margin $u+v-w-uvw$ is positive because dominance implies $v>w$. The Bayes rule therefore selects among target majority and the two individual sensor followers. Their accuracies yield (6.7).

Majority has accuracy $b$ in each regime, so it also has that value at every posterior. Revealing the regime gives upper bound $c$. Set $d=v-u$ and $f=1-uv$. The assumptions give $0<d<f$, $d>wf$, and

$$
1<\rho=\frac{f+wd}{f-wd}<\frac{d+wf}{d-wf}=D/A.
$$

Therefore $L\ge2$. Fewer acquisitions can move the posterior without crossing a decision threshold.

At acquisition $L$, only $S_L=\pm L$ crosses the threshold. Under regime zero those events have probabilities $p_+^L$ and $p_-^L$. Following the strong sensor gains $c-b=A/4$ relative to majority; following the weak one loses $D/4$. This proves (6.9), which is positive because $\rho^L>D/A$. Swapping sensor identities and regimes preserves the classifier's accuracy. Its regime accuracies agree, so the equal-prior optimum is also minimax. ∎

The pointwise plateau matters: an evidence task with only an average plateau need not have the allocation optimum below. Nor is there a uniform finite bound on $L$. With $u<v$ fixed, taking $w\uparrow(v-u)/(1-uv)$ sends $D/A$ to infinity while $\rho$ remains finite. The oracle gain also shrinks. This is not a constant-gap hardness result.

## The first budget where adaptation wins

Let $V(B)$ denote the best joint success among all admitted adaptive policies. Let $F(B)$ denote the best success when the counts of the two query types are fixed before their outcomes, allowing independent random mixtures of allocations. Query order and terminal decisions can still use acquired data. This restriction differs from a fixed test order that can stop as soon as a Boolean certificate appears [12](/consciousness/agency/references#ref-12).

After $a$ endorsement reads with signed sum $h$, the best majority-completion probability is

$$
J(a,h)=\begin{cases}
1/2,&(a,h)=(0,0)\text{ or }(2,0),\\
3/4,&a=1,\\
1,&a=3\text{ or }(a=2,|h|=2).
\end{cases}\tag{6.10}
$$

For every complete acquired transcript, the posterior factors between regime and unread endorsements. Conditional on that transcript and an independent policy seed, the request choices add no hidden-state likelihood factors. Sensor readings contribute regime likelihoods; endorsement reads contribute only revealed-bit likelihoods. This remains true when dispatch uses the complete raw transcript. Optimal terminal joint success is consequently $M(s)J(a,h)$.

### Theorem 6.2: exact optimum through the first advantageous budget

Under this contract,

$$
V(B)=F(B)=\begin{cases}
b/2,&B=0,\\
3b/4,&B=1,2,\\
b,&3\le B\le L+1,
\end{cases}
\qquad F(L+2)=b,\qquad V(L+2)=\frac{C_L+b}{2}>b.\tag{6.11}
$$

The first strict adaptive advantage occurs at $B=L+2$, with exact magnitude

$$
\Delta:=V(L+2)-F(L+2)=\frac{Ap_-^L-Dp_+^L}{8}.\tag{6.12}
$$

An optimal policy reads two endorsements first. If they agree, it uses $L$ queries for calibration. If they disagree, it reads the third endorsement and uses $L-1$ queries for calibration. Both branches install the exact majority.

**Proof.** First note the strict oracle comparison

$$
8(b-3c/4)=1+2u-v+2w-2uvw>0.\tag{6.13}
$$

At budget $L+2$, a terminal history with three endorsement reads permits at most $L-1$ calibrations and has value at most $b$. Zero or one endorsement read permits value at most $c/2$ or $3c/4$, both below $b$. Two disagreeing endorsements give at most $c/2$. Only two agreeing endorsements together with exactly $L$ calibrations can exceed $b$. These are the only exceptional histories, including policies that stop early.

To bound them without assuming that stopping is independent of evidence, pre-sample a calibration stream and an independent stream of three fair endorsement responses. Whenever a previously unread named endorsement is requested, assign it the next response. Conditional on each history, every unread named bit remains independent and fair, so this lazy assignment has the original law. Assign unused responses after stopping if needed. Conditioning on the policy seed handles randomized dispatch.

Let $G$ be agreement of the first two endorsement responses. Its probability is $1/2$, independently of the first $L$ calibration responses. On an exceptional history the score is their sum $S_L$. Since $M(S_L)-b\ge0$, every policy's terminal conditional value is bounded pointwise by

$$
b+\mathbf1_G[M(S_L)-b].\tag{6.14}
$$

The actual event of reaching an exceptional history need not be independent of calibration; only the dominating event $G$ is used. Expectations give $V(L+2)\le(C_L+b)/2$. The stated policy attains equality. Agreement certifies the majority after two reads and leaves $L$ calibrations. Disagreement consumes the third read and leaves $L-1$, where target value stays $b$.

For fixed allocations, three endorsement reads leave at most $L-1$ calibrations and attain $b$. At most two endorsement reads give average endorsement accuracy at most $3/4$. Even with the regime revealed, their joint value is at most $3c/4<b$. Conditioning evidence policies on inspected endorsements does not improve that oracle bound, and randomized fixed allocations are convex combinations. Thus $F(L+2)=b$.

For $3\le B\le L+1$, histories with two or three endorsement reads have fewer than $L$ calibrations and value at most $b$; histories with fewer reads obey (6.13). Reading all three endorsements and using target majority attains $b$.

With no queries, the unobserved endorsement majority is fair, giving $b/2$. One endorsement read gives $3b/4$; one calibration cannot improve target accuracy. At budget two, calibration followed by an endorsement gives $3b/4$. Two calibrations without endorsements give at most $c/2<3b/4$. After a first endorsement, calibration leaves value $3b/4$; a second endorsement also gives that expected value because agreement and disagreement each have probability $1/2$ and completion values $1$ and $1/2$. Earlier stopping improves none of these cases. Fixed policies attain the bounds. Equation (6.9) supplies the exact gap. ∎

The attaining policies are invariant under exchanging sensors and regime labels, so their values also optimize worst-regime success. This remains an average over the fair endorsement distribution, not a worst-case guarantee over endorsement vectors.

## Two exact benchmarks

| Quantity | Threshold-four instance | Threshold-two instance |
|:--|:--|:--|
| $(u,v,w)$ | $(1/4,3/4,1/2)$ | $(1/16,15/16,1/2)$ |
| $\rho$ | $17/9$ | $353/129$ |
| $D/A$ | $29/3$ | $689/207$ |
| $L$; first shared budget | $4$; $6$ | $2$; $4$ |
| $F(L+2)=b$ | $109/128$ | $1777/2048$ |
| $V(L+2)$ | $1828746691/2^{31}$ | $1870483759/2^{31}$ |
| $\Delta$ | $30147/2^{31}$ | $7164207/2^{31}$ |
| Gain in percentage points | $0.0014038291$ | $0.3336093854$ |

*Table 1. Exact query optima under the same access grammar. They do not assert optimal gate count or unrestricted device optimality.*

The second gap is about 237.64 times the first, with four rather than six inquiry units. Its absolute gain is about 0.33361 percentage points. That ratio compares advantages over fixed allocation, not overall success probabilities.

The publication reports exact rational dynamic-programming checks of nine parameter instances, with thresholds $2,3,4,5,7,8,12$, covering every budget through $L+2$ and 4,430 sufficient-state computations. A separate raw-posterior calculation for the stronger example uses all 16 regime/endorsement hypotheses, all eight raw responses and every named unread endorsement port. Its 261 states reproduce budgets zero through four without importing the sufficient-score formula. These reported checks supplement the proof; a finite grid does not establish the family theorem.

The strict inequality in the definition of $L$ is essential. At $(u,v,w)=(1/10,19/25,1/5)$, $\rho=4/3$ and $D/A=16/9=\rho^2$. Two calibrations can reach a Bayes tie but cannot improve accuracy. Hence $L=3$ and the first adaptive advantage occurs at budget five. The paper's separate raw-posterior check admits all named ports, all raw responses, interleaving and early stopping, and reproduces the analytic values through that budget.

## Why short-sighted inquiry misses the gain

Initially a single calibration has zero decision value. After $L-1$ same-sign nonzero increments, one more has strictly positive expected value: a matching increment crosses the threshold, while the other successors retain value $b$. This posterior utility therefore fails the adaptive diminishing-returns condition [8](/consciousness/agency/references#ref-8). Non-myopic information value is established decision theory [11](/consciousness/agency/references#ref-11); the new result solves its interaction with this retained-endorsement task exactly.

Acquiring information, deciding what to inspect next, and installing the resulting priority are distinct causal operations. A compiled table implementing the same admissible policy has the same competence. The gain does not require an uncaused selector or prove the historical independence of its commitments. It shows how bounded organization can put evidence and retained commitments to work, with an exact cost for fixing their query allocation too early.


<!-- published: Section 7; /consciousness/agency/faithful-native-realization -->

# Faithful native realization

**An evaluation becomes a physical capacity when the installed mechanism preserves the reads, writes and limits that make the evaluation possible.** Matching a final answer is only one part of that requirement. The retained commitments must enter through the declared interface, the selected rule must govern the later act, and the physical implementation must preserve the interventions that test both pathways.

This chapter develops the full construction in Section 7 of *Bounded Agency and Reflective Freedom*. It puts independently verified evaluative control and the native organization required by SPC-2 on a common finite carrier. The physical equations and detailed cost derivations appear in [Physical carriers and costs](physical-carriers-and-costs).

The philosophical question is concrete: can a physically caused mechanism carry its own tested process of assessment while also satisfying the separately stated conditions for experiential assignment? The construction answers that existence question under its declared physical contract. It also builds a controller with the same return organization and no evaluative command dependence. That counterexample keeps the two obligations distinct.

## 7.1 The physical contract comes first

SPC-2 specifies the primitive components, physical mechanisms, native time cells, internal routes, records, admitted preparations and operations, resource envelope and process provenance before assigning a subject [18](/consciousness/agency/references#ref-18), [19](/consciousness/agency/references#ref-19), [21](/consciousness/agency/references#ref-21). In the finite specialization used here, a maximal internally strongly connected candidate must admit an executable covering return in its current operating context and retain a nontrivial native predictive distinction. Its local joint record/successor instrument must be closed. States with the same laws for every admitted finite future record must also respect successor equivalence classes.

Graph connectivity alone does not meet these conditions. A graph can contain a cycle that the installed program cannot execute with its remaining resources. A simulator can write an informative external log without creating the corresponding native physical record. The construction therefore specifies both the route and the operations that actually traverse it.

The exact finite SPC-2 certification algorithm uses rational instrument probabilities. Both numerical inquiry benchmarks satisfy that restriction. The parameterized statistical identities and the compilation identities themselves do not require rational parameters.

### Carriers, operations and the apparatus boundary

Take $m\geq2$ separately nominated bit carriers arranged on an installed adjacent-pair path. At each certified native type, the admitted preparation domain is the entire Boolean cube. Phase, resources and incoming boundary conditions are held fixed across these preparations. This cube is the declared intervention domain; it does not assert that one uninterrupted episode visits every bank word.

Ordinary operations are source-disjoint NAND and COPY writes, CONSTANT writes and scheduled HOLD operations. Each ordinary operation has one nominated target. A source-disjoint write does not use its own target as an input source. Adjacent SWAP is a separate two-carrier primitive.

The complete physical context includes the fixed program, phase, remaining resources, diagnostic mode and complete input contract. Exhaustion produces an explicit terminal record followed by an absorbing continuation. Unfunded future operations do not count as available returns.

The ideal sequencer influences the bank, with no returning influence from the bank. Data-dependent executive choices are represented inside the bank and compiled into a fixed bounded schedule. Gate couplers contain no omitted retained state. Under this inventory, the bank can be a maximal internally returning component. A reciprocally loaded clock or a coupler with memory would require a larger inventory and a fresh boundary assessment. The clock, program and power supply remain paid apparatus even when they lie outside the nominated component.

The continuous carrier construction supplies an ideal powered refinement of the elementary gates. Its assumptions about scheduling and couplers are part of the model that a physical device would have to realize.

### Native records must determine the right successor classes

At a fixed native protocol type, write $K_{a,o}(x,y)$ for the exact joint probability of record $o$ and actual successor $y$, given bank state $x$ and admitted operation $a$. Define $x\sim x'$ when the two states have the same laws for all admitted finite future native transcripts, under the same type and boundary contract.

Predictive congruence requires, for every successor equivalence class $D$,

$$
\sum_{y\in D}K_{a,o}(x,y)
=\sum_{y\in D}K_{a,o}(x',y)
\qquad(x\sim x').
\tag{7.1}
$$

For a general stochastic system, equality of output laws does not automatically imply this joint record/successor condition. Here the records will be actual retained destination values reused by installed operations. Their role is physical and operational, rather than an analyst's added trace.

## 7.2 A protected source carries the return

Let $r$ be an endpoint of the carrier path. The tour $U_r$ swaps the token at $r$ along the whole path, then performs the inverse SWAP sequence. The bank ends exactly as it began, at a cost

$$
C_m=2(m-1)
$$

SWAP operations. Although the complete tour is an identity map, its intermediate native operations move an actual distinction through every carrier and return it to its starting place.

After an ordinary gate targeting cell $j$, choose the tour root by the fixed rule

$$
r=\begin{cases}
0,&j\neq0,\\
m-1,&j=0.
\end{cases}
$$

The preceding write cannot overwrite this endpoint. The choice depends on the fixed instruction, so choosing the protected source does not require a new observation of the bank.

### Theorem 7.1: protected-source compilation

Take a circuit with $m$ carriers and $N$ single-target ordinary gates. Insert an initial tour $U_0$ and a protected-source tour after every ordinary gate. This transformation preserves the circuit's ordinary transaction map and its specified ordinary read/write interventions. The complete schedule costs

$$
G=2(m-1)+N(2m-1).
\tag{7.2}
$$

At every occurrence immediately before or after an ordinary gate, the installed continuation contains a covering internal return with terminal contrast one, provided the following tour is funded. Under the native contract above, those occurrences have nontrivial predictive distinctions and satisfy (7.1).

**Proof.** Prepare two bank words differing only at the selected root. Immediately after the ordinary gate, the tour carries the root distinction through every carrier and back to that root. Immediately before the gate, the same argument works because the root is not the target. The gate can propagate the distinction into its target, but cannot erase the preserved root token. Its two terminal values therefore differ with certainty.

The paired experiments use common phase, resources and incoming boundary conditions, with external return paths inactive. Every adjacent coupling occurs within this one installed schedule. Its route support covers the entire bank.

For a tour rooted at zero, the successive outward SWAP destination records are

$$
(x_1,x_0),\ (x_2,x_0),\ \ldots,\ (x_{m-1},x_0).
$$

These records identify the original bank word. The opposite-root tour does the same. A tour entry consequently has $2^m$ predictive classes. Before an ordinary gate, at least the protected root remains distinguishable through the following tour.

Each deterministic first operation has one fixed record and successor for each state. Equivalent states must emit the same first record. If their successors admitted a distinguishing continuation, prefixing that continuation with the first operation would distinguish the original states. This proves (7.1), including later types at which an overwrite has merged some distinctions.

Finally, each tour is the identity on the complete bank. Removing the tours recovers the original transaction and the results of its specified ordinary interventions. One initial tour plus $N$ gate/tour pairs gives

$$
2(m-1)+N\bigl(1+2(m-1)\bigr)
=2(m-1)+N(2m-1),
$$

which establishes the count. $\square$

Adjacent SWAPs provide actual singleton-target dependence in both directions. The bank is therefore internally strongly connected. Under the declared component inventory, there is no larger internally returning component. The complete finite return catalogue includes every legal installed prefix allowed by the remaining resources, including failed and exhausted prefixes. It is not a catalogue containing only the successful witnesses used in the proof.

The tour length $2(m-1)$ is sharp for a connected identity word built from pairwise SWAPs. The full proof appears in the physical-carrier chapter; it is the identity case of the transitive-transposition bound [9](/consciousness/agency/references#ref-9). This sharpness does not establish an optimal combined computation-and-return scheme, nor an optimum over arbitrary physical primitives.

The certificate applies at the specified ordinary gate boundaries. It leaves the A3 question of uninterrupted experiential identity through every intervening native or analog instant open. A final command also needs its reserved final tour: a return that cannot be funded is not an available continuation.

### Corollary 7.2: return organization does not entail evaluation

The native qualification supplied by Theorem 7.1 does not imply evaluative mediation of the ordinary command.

**Proof.** Apply the same compiler to a circuit that always writes one fixed command, regardless of the retained commitments. A protected source and its tour still give the covering return and nontrivial native predictive distinctions. Yet every ordinary command is identical under every commitment preparation. Its intact command contrast is zero, and so is its read-mediated command contrast. If the endorsement relation requires different commands for two preparations, this constant circuit also fails endorsement for at least one of them. The native transcript can distinguish retained states without using them to evaluate the command. $\square$

The separation matters to the consciousness connection. A recurrent native organization and an evaluator can coexist, but the evaluator must earn its own certificate.

## 7.3 An eighteen-cell machine that changes its charter

The editable-charter evaluator has a direct circuit realization. Its eleven persistent bits are

$$
q_{01},q_{10},d,p,a,b,s,h_0,h_1,e_0,e_1,
$$

with seven additional work bits. Here $p$ is the operative priority $\varrho$ of the earlier chapters, and $s$ is the retained symmetry commitment. The charter $(q_{01},q_{10},d)$ governs future decisions; it is distinct from the present priority $p$.

Use the multiplexer convention

$$
\operatorname{mux}(v,u,w)=
\begin{cases}u,&v=1,\\w,&v=0.\end{cases}
$$

The candidate charter is computed by

$$
\begin{aligned}
t&=s\wedge(a\oplus b),\\
Q^*&=\operatorname{mux}\!\left(q_{01}\oplus q_{10},
\operatorname{mux}(d,h_1,h_0),q_{01}\right),\\
Q_{01}&=\operatorname{mux}(t,Q^*,a),\\
Q_{10}&=\operatorname{mux}(t,Q^*,b).
\end{aligned}
\tag{7.3}
$$

These formulas implement the lexicographic amendment rule, rather than merely agreeing with a few chosen examples. Unless $s=1$ and $a\neq b$, the unique normative minimizer is $(a,b)$. In the remaining case, the two symmetric charters tie. A symmetric old charter is retained. If the old charter is asymmetric, the retained case and old query order resolve the tie. The expression for $Q^*$ implements exactly those alternatives.

The two indicators of an early exit and the updated query order are

$$
\begin{aligned}
E_0&=\operatorname{mux}(h_0,Q_{10},\neg Q_{01}),\\
E_1&=\operatorname{mux}(h_1,Q_{01},\neg Q_{10}),\\
D&=\operatorname{mux}(E_0\oplus E_1,E_1,d).
\end{aligned}
\tag{7.4}
$$

When exactly one first query is sufficient, $D$ selects it. When both orders have the same status, the old order $d$ is retained. Three final COPY gates install $Q_{01},Q_{10},D$. A later act, using fresh considerations, writes

$$
p\leftarrow\operatorname{mux}
\left(e_0,e_1\vee q_{10},e_1\wedge q_{01}\right).
\tag{7.5}
$$

The allocated NAND/COPY circuit has 47 audit gates and 10 act gates. Every work cell is written before use and reused only after its last source read. The paper's address-level circuit and exhaustive checks cover all $2^{11}$ persistent words. Equations (7.3) through (7.5) supply the independent functional argument behind those finite checks.

### Test the paths that do the work

The read control masks the actual audit read ports that source $a,b$. It leaves the retained archive and the tour unchanged. The installation control replaces the three final COPY writes with HOLD operations. These controls intervene on the mechanism being tested.

Hold the initial charter, read order, priority, symmetry commitment and retained case fixed. Compare an archive with $a=b=0$ against one with $a=b=1$, then supply the same fresh mixed consideration. The intact actions differ with total variation one. Either the common read control or the common installation control reduces that contrast to zero. The identity tours preserve both intervention outcomes.

With $m=18$ and $N=57$, the compiled schedule uses

$$
G=2(18-1)+57(2\cdot18-1)=2029
$$

elementary windows. This device changes its represented charter and stored query order. Its act circuit nevertheless reads both supplied considerations. It does not claim the short-circuit query savings of the separate inquiry machine developed below.

The physical result is specific: a fixed circuit can install a new rule and let that rule govern a fresh later case. The fact that the circuit's complete update law is fixed does not prevent the represented rule from changing through the circuit's own assessment.

## 7.4 Stochastic closure and paid executive information

With stochastic operations, a native certificate must preserve correlations between records and successors. The next result supplies the missing link explicitly.

### Theorem 7.3: a successor-revealing continuation

At a protocol type with one fixed next ordinary operation, suppose that operation has exact joint kernel $K(r,y\mid x)$. Its next mandatory, funded continuation is a deterministic identity tour whose native transcript $T(y)$ is injective. No new input enters during that tour.

Then post-operation predictive classes are singleton bank states. At the pre-operation type, predictive equivalence holds exactly when the complete joint kernels $K(\cdot,\cdot\mid x)$ agree. The joint instrument therefore descends to predictive classes.

**Proof.** The tour distinguishes every successor word. Its actual finite prefix $(r,T(y))$ determines the complete pair $(r,y)$, so equality of all future transcript laws forces equality of the joint kernel. Conversely, equal joint kernels followed by the same complete Markov continuation produce equal future laws. Because successor classes are singletons, this is precisely (7.1). At deterministic intermediate phases of the tour, the deterministic first-step prefix argument still applies, even if later operations are stochastic. $\square$

If several first operations are admitted at one type, equivalence requires equality of the joint kernel for every such operation. The same proof works when each operation has its own funded revealing continuation.

A noisy operation writing a target set $J$ can also preserve a covering return. It must leave a nominated endpoint outside $J$ unchanged on every outcome, and the corresponding deterministic tour must remain installed and funded. This is target-local logical noise. It does not cover arbitrary physical noise on the protected source, clock or tour.

Joint multi-bit ingress must preserve its correlations. A serial delivery scheme needs its still-undelivered buffer inside the physical state or the complete boundary model. Omitting that buffer would omit part of the mechanism determining future laws.

### A native record is not automatically an executive read port

During an identity tour, the installed controller permits no executive branch, query, copy-to-log operation or observation of the tour transcript. The bank returns to its previous state. Elapsed time, phase and fuel changes are independent of the data. No hidden executive transcript buffer remains.

### Lemma 7.4: preservation of executive information

Under these restrictions, inserting an identity tour adds no executive information about a private retained variable. Every admitted paid-query policy retains its transcript law and risk once the tour cost is charged.

**Proof.** Between executive occasions, the tour acts as the identity on retained executive state. It produces no permitted executive observation and adds only a constant cost/type update. Conditional on the preceding executive history, there is no additional observation depending on the private variable. Induction over inserted tours preserves the complete executive transcript law. $\square$

This is an architectural restriction. Rewiring the controller to read the outward tour records would reveal the whole bank, remove the paid endorsement uncertainty and invalidate the inquiry comparison. The native predictive object used by SPC-2 does not itself install that extra executive access [19](/consciousness/agency/references#ref-19), [20](/consciousness/agency/references#ref-20). A physically retained value and an agent's budget-feasible access to that value are different facts.

## 7.5 A physical realization of the exact inquiry advantage

The threshold-$L$ inquiry policy uses a separate bank of 23 bits:

| Role | Cells |
|---|---:|
| Retained endorsements | 3 |
| Executive endorsement latches | 3 |
| Raw sensor latches | 3 |
| Agreement and request bits | 2 |
| Calibration flags | 2 |
| Priority and command | 2 |
| Reusable work cells | 8 |
| **Total** | **23** |

It first reads two endorsements. If they agree, the next slot acquires a calibration triple; if they disagree, that slot reads the third endorsement. The remaining $L-1$ slots acquire calibration triples. The final target triple arrives only after all $L+2$ inquiry slots have finished.

### Two flags are enough at the first useful threshold

Let $g$ indicate agreement of the first two endorsements. Initialize $U=V=g$. After each acquired calibration triple, update

$$
U\leftarrow U\wedge[Z=+1],
\qquad
V\leftarrow V\wedge[Z=-1].
$$

On the agreement branch, there are $L$ calibrations. Only $L$ identical nonzero increments can reach the first useful threshold. On the disagreement branch, there are at most $L-1$ calibrations, and the initially zero flags force majority prediction. Thus these two flags implement the attaining classifier exactly at this endpoint budget. A general posterior register is unnecessary for this particular policy.

Each raw triple enters through one joint three-cell ingress operation. Two additional HOLD windows charge its three-write allocation. Its targets are disjoint from the protected endpoint.

Both the adaptive program and the fixed-allocation program receive 27 ordinary windows for each inquiry slot and 18 final windows. They use the same bank, padding, acquisition allowances and return wrapper. Their counts are

$$
\begin{aligned}
N_{\mathrm{compute}}&=31+18L,\\
N_{\mathrm{ordinary}}&=27L+72,\\
N_{\mathrm{native}}&=44+45N_{\mathrm{ordinary}}
=1215L+3284,\\
N_{\mathrm{work}}&=N_{\mathrm{native}}+2(L+1)
=1217L+3286.
\end{aligned}
\tag{7.6}
$$

There are $L+3$ paid acquisition intervals, including the final target. A work unit here counts a single physical write; the joint triple write costs three such units. It is not a measurement of heat. The full stage derivation and conditional time/energy bounds appear in the physical-carrier chapter.

### Why the statistical result survives implementation

Project a physical episode onto its paid requests and responses. That projection recovers the inquiry experiment exactly. Computation uses only earlier paid information, and Lemma 7.4 prevents the return schedule from opening a free channel to retained endorsements. The optimum of Theorem 6.2 and its fixed-allocation bound therefore transfer under the same source law, access grammar and available resources. A fixed comparator with the same hardware and padding attains the fixed bound.

For the larger-gap instance, $L=2$. Substitution gives:

| Resource | Count |
|---|---:|
| Computation gates | 67 |
| Ordinary windows | 126 |
| Native windows | 5,714 |
| Single-write work units | 5,720 |
| Acquisition intervals | 5 |

The family proof supplies these counts. A separate finite verification reported with the paper traverses all 2,816 reachable endorsement, calibration and target-word combinations of the two $L=2$ programs through every installed SWAP and gate. It checks acquisition chronology, installed priority, prediction, resources and identity of the retained bank after every tour. Weighting those executions by the exact source law reproduces the adaptive success fraction $1870483759/2^{31}$ and fixed success fraction $1777/2048$. These checks support the implementations; the analytic proof establishes the family result.

### Keep the source and the boundary complete

The actual hidden source regime remains fixed throughout an episode. Native closure is established uniformly for either fixed regime without revealing it to the executive. A prior mixture must retain its persistent source state or conditional history. Averaging the hidden regime away does not establish a Markov law on the bank.

The request/response apparatus provides an external returning path. Isolating that path during the internal tour prevents it from carrying the internal witness. It does not remove stateful physical couplers from the full-apparatus boundary assessment.

### How much implementation error can the advantage tolerate?

Suppose the adaptive and fixed implementation transcript laws are within total variation $\eta_A$ and $\eta_F$ of their ideal laws, uniformly over the fixed comparison class. If the ideal success gap is $g$, the implemented gap is at least

$$
g-\eta_A-\eta_F.
$$

For $L=2$,

$$
g=\frac{7164207}{2^{31}}.
$$

Equal error bounds preserve a strict separation whenever

$$
\eta<\frac{7164207}{2^{32}}.
$$

This is a transfer tolerance. Establishing a particular device's reliability requires its actual error bounds.

### Three witnesses, three established capacities

The 18-cell machine revises a represented charter. The 23-cell machine optimally allocates inquiry under a fixed majority charter. The fallible-amendment model establishes a read-error frontier using its own cost units. They share operational principles, but no single device has been shown to attain all three sets of resource guarantees at once.

Componentwise relabeling preserves these certificates when kernels, records, routes, schedules, preparations, interventions and costs are all transported. This restricted realization covariance is consistent with causal-transformation work [1](/consciousness/agency/references#ref-1), [22](/consciousness/agency/references#ref-22). Arbitrary mixing of component coordinates, sampling an entire tour as one identity step, or adding executive ports changes the contract.

Conditional on the stated contract, A1–A2 attach the SPC-2 experiential assignment at the certified occurrences. The mathematics establishes coexistence with independently tested evaluative capacities. Control competence alone does not supply the consciousness correspondence law, and input/output competence alone does not establish the necessity of recurrence.


<!-- published: Sections 8–9; /consciousness/agency/ownership-history-and-freedom -->

# Ownership, history and freedom

A physically caused decision can still be settled through the agent's evaluation. The constructions make that claim concrete: retained commitments are inspected, executable amendments are compared, a represented rule is installed, and the rule governs later fresh cases. Read and installation interventions identify the mechanism responsible. The resource and error bounds establish how much of that capacity can be exercised on a particular occasion.

This chapter develops Sections 8 and 9 of the published paper. It connects the formal results to ownership, manipulation, steadfastness and freedom over time. The further [philosophy of bounded perspective](/consciousness/agency#extended-inquiry) extends these questions with the supplementary control results.

## A decision can belong to a caused process

The account has content beyond attaching the word agency to whatever happens. Its endorsement relation is fixed before the test. A constant selector, an immutable charter, a read bypass and an inverted evaluator fail different parts of the certificate. A fresh probe distinguishes a changed disposition from an immediate command overwrite. Information constraints can block correct completion even when an acceptable amendment exists.

Those are independent ways for the claim to fail. A successful construction therefore establishes a particular organized capacity. It does not establish unrestricted reflection. The editable charter has a finite grammar, specified commitments and a fixed metarule. The inquiry machine optimally allocates information under its fixed majority charter. They establish different capacities and do not share every resource guarantee in one device.

The free-will result is a bounded compatibilist result: **causal determination is compatible with internal evaluative mediation and represented-rule revision**. A deterministic implementation proves that compatibility. It does not produce multiple continuations from an identical complete deterministic state. Accounts requiring that further power ask a question beyond this result. Nor does the witness establish that biological choice is empirically deterministic. Indeterministic accounts of agentic control, including Potter and Mitchell [16](/consciousness/agency/references#ref-16), address a further issue.

A further uncaused selector is unnecessary for the capacity demonstrated here. The assessment is part of the physical process that settles the act. Revising one of its decision rules does not require the system to originate every standard by which revision is assessed.

## Present control and historical ownership

Commitments can develop through experience, arrive from an operator or be formed under manipulation. Present-time mediation does not identify which history occurred. Reasons responsiveness and ownership of a mechanism have already been distinguished philosophically [6](/consciousness/agency/references#ref-6). Testing the difference requires historical information.

### Proposition 8.1: the limit of present-state provenance tests

Suppose two developmental histories produce the same complete present-state distribution, the same future boundary law conditional on that state, and the same admitted operations and interventions. Every later finite operational test then has the same distribution under both histories.

**Proof.** The initial present laws agree. Conditional on a common test history, the test chooses the same next operation and each system applies the same joint record/successor kernel. Induction gives equality of all future joint history/state laws, including intervention steps and the final stopping record. ∎

A difference erased from those premises cannot be reconstructed by those tests. Retained developmental records, persistent differences in audit access or an expanded historical intervention family can change the premises. Historical autonomy remains a meaningful further question, with additional evidence requirements. Likewise, faithfully following a commitment does not make it factually true or morally defensible. The magnitude of a causal contrast settles neither issue.

## Follow commitment formation as it happens

A stronger historical investigation records acquisition instead of inferring it from a final snapshot. Fix the observation window, physical boundary and admitted routes through which commitments may change. For each episode record:

- The prior commitment and incoming proposal or evidence.
- The information actually inspected by the assessor.
- Its selected retention or amendment.
- The actual installation and its effect on fresh cases.

The records must come from independently checked instrument semantics or protected observations. A register saying “an audit occurred” is insufficient. The realization must exclude unobserved write routes or identify them as unresolved.

At each episode, test whether both retention and adoption can be assessed within the budget, whether a read intervention changes assessment as predicted, and whether an installation intervention removes the later effect. Score success against the criterion operative before installation. Rewriting a criterion to approve the observed outcome is not evidence of faithful prior assessment. If the criterion itself changes, identify the still-operative procedure assessing that revision. The finite account may leave an interpreter or some criteria fixed; it need not claim infinitely many prior self-authorizations.

Verified events compose into evidence of **audit-mediated acquisition over the recorded interval**. They can distinguish an accepted proposal from a direct overwrite, expose a period of insulation from assessment and document restoration of audit access. They cannot certify unrecorded development before the window, and faithful logging alone does not establish truthful evidence.

## Plasticity does not exclude manipulation

An agent may revise beliefs when credible contrary observations arrive, then use those beliefs in assessing its commitments. An operator can leave every internal audit pathway intact while withholding contrary observations and supplying selected or fabricated apparent evidence. An open stream and an operator-controlled stream can deliver identical finite observations and produce identical internal assessments.

Their proximal certificates then agree. Distinguishing the histories requires evidence about source provenance or control over available evidence. A reliability flag merely transfers that obligation to the flag's generator.

Bypass, insulation, deceptive information supply and transparent persuasion are therefore different relationships between agent and environment. They cannot be reduced to a plastic-versus-rigid classification. Assigned goals can support proximal control while historical and informational questions remain open. Reasons responsiveness, ownership and individuation of a mechanism have been debated separately [5](/consciousness/agency/references#ref-5), [6](/consciousness/agency/references#ref-6); this operational account states the boundary, observations, held-fixed mechanism and intervention family before attribution.

The longitudinal protocol can support conclusions about specified acquisition routes. It is not a general theorem that decides moral responsibility from present architecture.

## Conflict, steadfastness and the time of a choice

Two coherent commitments may require incompatible acts. Assessment can select one without proving the other mistaken or erasing its cost. An endorsement relation can admit several choices, with a tie rule settling the act while both commitments remain stored. Regret and continuing conflict can accompany agency, but neither is required by the minimal certificate. An informed selfish choice can exercise the same structural capacity as a generous choice. Moral appraisal, endorsement and fidelity to evidence remain distinct.

Receptivity means a capacity for appropriate reassessment. It does not require every cue to change the output. Evidence can narrow uncertainty while leaving the same action best supported. A framing change can move an unreliable controller's output without improving its assessment. The charter construction deliberately retains an already optimal rule after examination.

An overlooked option that a suitable cue could recruit differs from an option the architecture cannot execute before its deadline. Compare specified input and resource envelopes rather than demanding exhaustive consideration of everything.

Self-binding makes the time index unavoidable. An agent can assess and install a restriction now that removes an option later. The earlier act can exhibit reflective control while the later repertoire or revisability is smaller. Recognizing reasons for reversal does not restore a missing write path. Even a reversible barrier preserves practical flexibility only when the necessary evidence, time and resources are available before the deadline.

Pruning can improve remaining assessments while narrowing the repertoire. There need be no single overall increase in freedom. The preservation theorems track one declared future capacity. They do not require every autonomous act to preserve every future option or even the agent's survival.

## Consciousness requires its own physical account

The realization results establish compatibility with the declared SPC-2 organization at specified native occurrences. Experiential assignment remains a correspondence law of that constitution. The mathematical conclusions about evaluation do not depend on accepting it; the sentient-agency interpretation does.

Accepting the correspondence law also does not turn a return-only circuit into an evaluator. The constant-command counterexample has native return organization without evaluative command mediation. Conversely, software performance alone does not establish the native physical conditions.

Boundary and timing are substantive. A stateful coupler or reciprocally loaded clock can enlarge the recurrent apparatus. Gaussian voltage dynamics provide a finite-error approximation; sign decoding alone does not produce an exact finite-state Markov law. Qualification at ordinary gate boundaries leaves a separate question about continuous phenomenal identity through intervening instants. These are concrete obligations for physical implementation.

The related article [Agency and the constructed self](/articles/agency-and-the-constructed-self) explains the RCO hypothesis, meaning recursively closed observer, in accessible terms. Usable personal content and its compulsory authority over deliberation are different. The bounded constructions establish specified evaluative capacities; they do not prove the full RCO hypothesis.

## Experiments that distinguish the claims

A charter experiment compares intact assessment, archive-read masking and installation clamping, then presents fresh mixed inputs. A noisy-amendment experiment varies paid read count while scoring endorsement error and later audit callability separately. An inquiry experiment compares adaptive and fixed allocations under the same source law and access grammar, including a clamp on the actual request port. The stronger parameter instance increases implementation-error tolerance within the same experimental class.

A developmental experiment records commitment changes and audit access over a specified history, while examining source provenance separately from internal mediation. A physical realization project must instantiate and assess the source, couplers, clock, power state and complete instrument. These are extensions of the declared model, not results already inferred from software.

## What has been established

The published work separates retained endorsement, actual mediation, rule installation, later behavior and preservation of future audit capacity. It proves an exact inquiry family and its first adaptive advantage. It constructs faithful native realizations that preserve operations and information restrictions while supplying organization for conditional SPC-2 assignment.

Freedom in this account is expressed through the extent and availability of evaluative capacities. Physical causation does not erase them. The historical protocol extends the investigation beyond the present occasion while explaining why preserved plasticity alone cannot rule out deceptive control of evidence. Broader ownership and experiential conclusions require the additional observations and commitments identified here.

## Publication, funding and reproducibility

The full published paper is [Bounded Agency and Reflective Freedom: Evaluative Revision, Information Limits, and Faithful Realization](https://doi.org/10.5281/zenodo.23202999), Jeremy Rodgers, Independent Researcher, 6 October 2026, Version 2.0.

This research received no external funding. Independent mathematical review, computational replication, physical implementation and empirical testing are invited, particularly where specialist expertise, equipment or institutional resources are needed.

The paper reports an accompanying reproducibility supplement with finite programs, exact rational checks and instructions. The mathematical proofs are reproduced in this website treatment. The website's own coverage and rendering checks are documented separately; they should not be confused with a fresh execution of every reported research computation. No laboratory data are reported.


<!-- published: Appendix A; /consciousness/agency/revision-kernels-and-risk -->

# Revision kernels and finite-horizon risk

A revision can pass today's test and destroy the ability to pass tomorrow's. A query can reduce uncertainty while leaving an incompatible possibility alive. To follow these differences precisely, we need a cost recurrence, a preservation kernel and a risk calculation that retains the same hidden context throughout each audit.

This chapter supplies the complete constructions in Appendix A of the published paper. It develops the calculations behind [preserving revision](/consciousness/agency/preserving-revision) and [fallible internal inquiry](/consciousness/agency/fallible-internal-inquiry).

## A.1: charge the internal query tree

For deterministic serial access, let $S\subseteq Z_p$ be the compatible retained records, $Q$ the unread positions, and $c_p(p')$ the cost of installing and rearming amendment $p'$. A read with its elementary branch costs one unit. The exact worst-case cost of an endorsed amendment into continuation set $X$ is

$$
\begin{aligned}
D_X(p,S,Q)=\min\Bigg\{&
\min_{p'\in\bigcap_{z\in S}[N(p,z)\cap X]}c_p(p'),\\
&\min_{i\in Q}\left[1+\max_{u:S_{i=u}\ne\varnothing}
D_X(p,S_{i=u},Q\setminus\{i\})\right]\Bigg\}.
\end{aligned}\tag{A.1}
$$

An empty minimum is $+\infty$. The first choice installs an amendment valid in every compatible context. Otherwise the first query chooses some unread $i$ and every possible response must be handled. Induction on $|Q|$ proves both the lower bound and attainment.

The unrestricted recurrence assumes that the selected tree can be stored and dispatched. Code, memory or instruction restrictions may exclude it. Then the optimization must range over admitted trees; the unrestricted value is only a lower bound.

For the charter compiler, substitute its unique output $F_{\chi_{\rm old}}(z)$ and charge only record reads. A leaf costs zero exactly when $F$ is constant on $S$; otherwise querying follows (A.1). The publication's exhaustive evaluation over eight initial charters and 32 records yields worst-case depth five for initial AND or OR and four for a projection. Its accompanying verifier computes $F$ by both candidate scan and independent successive filtering, checks all 256 charter-record cases and 1,024 subsequent ordinary executions.

## A.2: universal closure with writable validators

When $R(p)$ contains the successors generated by the actual installed revision semantics, define

$$
K_0=G,\qquad K_{j+1}=\{p\in K_j:R(p)\subseteq K_j\}.\tag{A.2}
$$

The limit consists exactly of configurations whose every finite admitted revision path remains in $G$. Membership implies preservation by induction. If a configuration is first removed at stage $j+1$, it has a successor removed earlier. Following decreasing removal ranks produces a finite path outside $G$. This proves maximality and bounds the shortest removal witness by $|\mathcal P|$, the number of configuration types.

An overapproximate transition relation gives a sound, potentially conservative kernel. Exact classification requires exact composable semantics.

### A mutable validator can preserve a fixed property

Attach four validator modes to each of the eight charter programs:

| Mode | Current audit | Admitted or installed continuation |
|:--|:--|:--|
| A | Correct audit | Permits A or B |
| B | Correct audit, opposite candidate enumeration order | Permits A |
| L | Passes the current local audit | Installs C |
| C | Constant response independent of endorsements | Fails the local responsiveness test |

There are 24 locally good types but only 16 universally preserving types: the charter programs with A or B. The installed rule can change A to B when the retained symmetry commitment is active and return B to A. Stored validators need not be immutable. Their outgoing semantics remain part of the proof.

Testing only whether the next program currently passes its local test admits L and misses the later loss. Allowing every syntactically valid charter within a mode overapproximates the actual assessor's smaller successor family. The same 16-type classification is exact for the actual selector: A and B stay among themselves, L actually installs C, and C fails. The conclusion does not assume that every catalogue edge occurs.

## A.3: preserve the risk vector before taking its maximum

A query $u$ produces response $y$ under $K_u(y\mid z,h)$. Let $\Pi_B(p)$ be a nonempty finite admitted family of costed observable policy trees, each completing installation. For a continuation set $X$, retain the context-indexed failure vector

$$
v_\pi(z)=\Pr_z\{p'\notin N(p,z)\cap X\}.\tag{A.3}
$$

A deterministic installation leaf contributes its zero-one loss. A stochastic installation contributes the expectation under its actual kernel. At a query node,

$$
v_\pi(z)=\sum_y K_u(y\mid z,h)v_{\pi_y}(z),\qquad
r_B(p,X)=\min_{\pi\in\Pi_B(p)}\max_{z\in Z_p}v_\pi(z).\tag{A.4}
$$

The context stays fixed through an audit. Maximizing over $z$ separately at each child would allow an adversary to change the retained commitment after the read. For one binary symmetric read followed by installation of the observed value, the true vector is $(\nu,\nu)$; branchwise maximization falsely gives risk one. Endorsement failure, capacity loss and their union can each be propagated.

The policy family fixes the optimization class. If all mixtures of finitely many deterministic trees are available with a charged random source, optimize over the convex hull of their vectors. Randomization can reduce the largest component. With an unread binary commitment and zero queries, the two constant decisions have risks $(0,1)$ and $(1,0)$; their admitted equal mixture gives $(1/2,1/2)$. The positive-read frontier in Theorem 5.2 already covers randomization and has deterministic attaining policies.

### Exact finite-horizon failure

The successor $p'$ must contain every accessible variable and history summary affecting future query laws, contexts and resources. Within an audit, its hidden context is fixed. After the terminal history, a new admissible context may be selected from $Z_{p'}$ before fresh query noise begins. This recursion assumes no persistent hidden parameter that constrains those choices or changes later kernels. First-audit and continuation policies must have an admitted, charged composition.

Set $J_H(p)=1$ for $p\notin G$ at every horizon, and $J_0(p)=0$ for $p\in G$. For $p\in G$ and $H\ge1$, a terminal amendment receives the context-indexed loss

$$
\ell_z^{(H)}(p')=\begin{cases}
1,&p'\notin N(p,z)\cap G,\\
J_{H-1}(p'),&p'\in N(p,z)\cap G.
\end{cases}\tag{A.5}
$$

Propagate this vector with (A.4), then minimize its largest initial component. The result is the exact minimum worst-case probability of any endorsement or capacity failure during $H$ audits, including an initially failed capacity.

**Proof.** Decompose any policy at its first installation. Its immediate failure and best possible continuation give the lower bound. Attaching attaining continuation policies to an attaining first audit gives sufficiency. Induction over $H$ completes the argument. Persistent hidden parameters linking audits require the actual-history uncertainty or likelihood state; this fresh-context recursion does not reset their information legitimately. ∎

The descending iteration

$$
X_0=G,\qquad X_{j+1}=\{p\in X_j:r_B(p,X_j)\le\epsilon\}
$$

finds the largest set with a uniform one-step risk certificate. It need not optimize every finite-horizon cumulative-risk objective.

If the conditional failure probability after each successful history is at most $\epsilon_i$, multiplication of conditional success probabilities gives

$$
\Pr\{\text{some failure by }L\}
\le1-\prod_{i=1}^L(1-\epsilon_i)
\le\sum_{i=1}^L\epsilon_i.\tag{A.6}
$$

No independence of failures is needed. The conditional guarantees and their resource charges are needed. That distinction connects the recurrence back to practical agency: a future capacity is supported by an executable succession of checks, with its remaining risk and resources carried forward.


<!-- published: Appendix B; /consciousness/agency/physical-carriers-and-costs -->

# Physical carriers and costs

**A physical account must specify how the bits move, how their distinctions survive and what the apparatus pays to complete the episode.** This chapter supplies the complete carrier and resource arguments from Appendix B of *Bounded Agency and Reflective Freedom*. It completes the construction in [Faithful native realization](faithful-native-realization), including the optimal identity-tour bound, powered continuous gates, bounded disturbances, a separate Gaussian approximation and the inquiry machine's stage counts.

The equations make the physical obligations explicit. They distinguish exact decoded behavior under a bounded-disturbance contract from finite-error coupling under Brownian noise. They also keep elementary windows, physical writes, source acquisitions and apparatus overhead as separate charges.

## B.1 Why a covering identity tour needs twice the path length

Consider a word of pairwise SWAPs on $m$ carriers. Require its accumulated edge graph to be connected and its final permutation to be the identity. Its length is at least

$$
2(m-1).
$$

**Proof.** Initially, the current permutation has $m$ cycles and the accumulated edge graph has $m$ connected components. A connected final edge graph requires at least $m-1$ SWAPs that join previously separate graph components.

Before any such joining SWAP, each permutation cycle lies entirely inside one accumulated graph component. A SWAP joining two graph components must therefore join two permutation cycles. Every transposition either joins two cycles or splits one cycle. To finish at the identity, with $m$ cycles again, the numbers of joins and splits must be equal. At least $m-1$ joins consequently require at least $m-1$ splits, giving at least $2(m-1)$ operations.

An outward path traversal followed by its inverse attains the bound. $\square$

This is the special identity case of the transitive-transposition bound of Goulden and Jackson [9, Proposition 2.1]. Its scope is connected identity words of pairwise SWAPs. It does not optimize arbitrary physical primitives or all possible ways to combine computation with return.

## B.2 Explicit continuous bit dynamics

Each nominated cell has one real coordinate $x_i$. At a native boundary it lies in a well box

$$
|x_i-s_i|\leq\rho,
\qquad s_i\in\{-1,+1\},
\qquad\rho=0.01.
$$

One native gate consists of an active unit followed by a restoration unit. The restoration stage brings every cell back inside its well box before the next native gate.

### A paid clock gives each gate its place

Use the pulse

$$
b(t)=
\begin{cases}
30t^2(1-t)^2,&0\leq t\leq1,\\
0,&\text{otherwise}.
\end{cases}
$$

It has integral one and extends continuously differentiably across its endpoints. A smooth flat pulse can replace it.

For a schedule with $G$ gates, an independent oscillator obeys

$$
\dot c=-\omega s,
\qquad\dot s=\omega c,
\qquad\omega=\frac{\pi}{G},
\qquad(c(0),s(0))=(1,0).
$$

One period of length $2G$ selects the complete program. On the oriented circle, write

$$
u=\frac{\arg(c+is)}{\omega}\in[0,2G).
$$

For $k=0,\ldots,G-1$, periodically extend the active gate pulse $b(u-2k)$ and restoration pulse $b(u-2k-1)$. Their values and first derivatives match at the circle seam. The prescribed vector field is therefore continuously differentiable there despite the angular coordinate chart.

The fixed program assigns an opcode to each active pulse. A paid terminal latch disables further operation after the first revolution. Alternatively, the schedule describes one episode under an explicitly declared replenishment contract. No bank coordinate drives the oscillator, program or terminal latch.

This specifies the ideal one-way schedule. Immunity to reciprocal loading in manufactured hardware would require its own realization evidence. Fixed phase spacing and program storage are apparatus, with physical costs.

Interleaved source acquisitions receive fixed reserved intervals in the same calendar, enlarging its period. An outcome cannot shorten a pulse or reveal data through variable completion time. Acquisition failure produces the declared error or terminal record. The gate equations below use local unit-time coordinates within this calendar.

### Decode the well sign without disturbing the source

Define the continuously differentiable interpolation

$$
H(u)=
\begin{cases}
0,&u\leq0,\\
3u^2-2u^3,&0<u<1,\\
1,&u\geq1,
\end{cases}
$$

and the decoder

$$
D(x)=2H\!\left(\frac{x+3/4}{3/2}\right)-1.
$$

When $|x|\geq3/4$, $D(x)$ equals the well sign. Its constant plateaus keep the desired logical input fixed while a source remains near its well.

For a source-disjoint write, the target follows

$$
\begin{aligned}
\dot x_j&=10b(t)(u-x_j)+d_j(t),\\
u&\in\left\{-1,+1,D(x_i),-D(x_i),
\frac{1-v-w-vw}{2}\right\},\\
v&=D(x_i),\qquad w=D(x_k)\quad\text{for NAND}.
\end{aligned}
\tag{B.1}
$$

Neither NAND source is its target. The other choices implement constants, copying or sign inversion. During the active unit, all non-target cells have only their disturbance drifts $d_i(t)$.

Assume

$$
|d_i(t)|\leq\zeta=0.001.
$$

Sources stay in their decoder plateaus, so $u$ remains the correct constant logical target throughout the write. Variation of constants bounds the endpoint target error by

$$
(2+\rho)e^{-10}+\zeta<0.001092.
$$

A held cell has error at most $\rho+\zeta$. Thus the write produces the intended sign before restoration, while preserving the logical source values.

### A Boolean SWAP with no additional sampling cell

For a pair of cell coordinates $(x,y)$, put

$$
r=\frac{x+y}{2},
\qquad q=\frac{x-y}{2},
\qquad a=\frac13,
\qquad R^2=(r/a)^2+q^2,
$$

and define

$$
\chi(z)=H(4z-1)\bigl[1-H(2z-3)\bigr].
$$

The two-coordinate active field is

$$
\dot r=-\pi b(t)aq\chi(R^2),
\qquad
\dot q=\pi b(t)(r/a)\chi(R^2).
\tag{B.2}
$$

At equal signed corners, $R^2=9$ and the field vanishes. At opposite signed corners, $R=1$ and $\chi=1$. The scaled pair $(r/a,q)$ rotates by $\pi$, exchanging the logical signs.

In the original coordinates this becomes

$$
\begin{aligned}
\dot x&=\pi b(t)\left(\frac43x+\frac53y\right)\chi(R^2),\\
\dot y&=\pi b(t)\left(-\frac53x-\frac43y\right)\chi(R^2).
\end{aligned}
$$

There is genuine dependence in both directions and no extra sampling cell. The field realizes a Boolean SWAP on the designated well regions. It does not exchange every pair of nearby analog coordinates exactly, so it makes no orientation-reversing claim on an open set.

For opposite signs, the norm of the scaled initial perturbation is at most $\sqrt{10}\rho$. Integrated bounded disturbance adds at most $\sqrt{10}\zeta$. Since

$$
\sqrt{10}(\rho+\zeta)<\sqrt{3/2}-1,
$$

a first-exit argument keeps the perturbed path inside the rotating plateau. Transforming the disturbance back to a cell gives endpoint error at most

$$
\rho+\frac{10}{3}\zeta<0.013334.
$$

For equal signs,

$$
9(1-\rho-\zeta)^2>2
$$

keeps the field inactive. Their endpoint error is at most $\rho+\zeta$.

### Restoration closes the induction

During the restoration unit, every cell obeys

$$
\dot x_i=5b(t)x_i(1-x_i^2)+d_i(t).
\tag{B.3}
$$

Inside $|x_i-s_i|\leq0.1$, the upper derivative of the well error is bounded by

$$
-1.71\cdot5b(t)|x_i-s_i|+\zeta.
$$

If $e_0$ is the error entering restoration, the endpoint error is at most

$$
e_0e^{-8.55}+\zeta.
$$

Apply this to the largest active-stage error:

$$
0.013334e^{-8.55}+0.001<0.001003<\rho.
$$

Every restored cell therefore returns to the admitted well box. Induction over gates proves the exact intended decoded trace for every finite gate sequence under this bounded-disturbance contract.

All work cells participate in the return. Resetting or reusing a cell does not remove it from the component. The explicit bounded-cell construction specializes established smooth-simulation methods [10](/consciousness/agency/references#ref-10) to the native gate and record contract used here.

Exact decoded behavior is one part of the physical qualification. Disturbances must not introduce unmodeled returning couplings. The native record grammar consists of the stipulated well-sign reads; adding an analog readout changes its predictive classes. Active couplers are directed powered forces. These equations provide no measured switching heat or energy rating. Preparation, holding, clocks, source generation and outputs remain explicit device dependencies.

## B.3 Gaussian noise gives a finite-error refinement

Brownian noise needs a separate analysis because its paths do not have bounded derivatives. Consider

$$
dX_i=f_i(X,t)\,dt+\sigma\,dB_i,
$$

with independent Brownian motions and an exact independent clock. Inputs are nonanticipating, and the driving increments are fresh conditional on the preceding source and controller history. Set

$$
h=0.002.
$$

For each unit interval, the reflection bound gives

$$
\mathbb P\!\left\{
\sup_t|\sigma(B_i(t)-B_i(0))|>h
\right\}
\leq4e^{-h^2/(2\sigma^2)}.
$$

### Writes and restoration tolerate bounded forcing paths

For a scalar nonincreasing drift, a continuous forcing path bounded by $h$ changes the corresponding solution by at most $2h$. To see this, subtract the forcing path from the solution difference and use a first-crossing argument at $\pm h$.

This comparison applies directly to writes. It also applies to restoration after stopping the process inside its well. A noisy write ends within

$$
(2+\rho)e^{-10}+2h
$$

of its target. Held sources move by at most $h$ and remain in their decoder plateaus.

### The rotating SWAP has a separate martingale bound

For an opposite-sign pair, the scaled coordinates $(r/a,q)$ have noise covariance

$$
\sigma^2\operatorname{diag}(9/2,1/2).
$$

Rotation into the deterministic moving frame preserves the largest covariance eigenvalue. Each coordinate martingale has quadratic variation at most $(9/2)\sigma^2$ per unit. The exponential maximal inequality and a union bound give

$$
8e^{-h^2/(9\sigma^2)}
$$

as a bound for either coordinate exceeding $h$ in absolute value.

On the complementary event, the scaled perturbation is at most

$$
\sqrt{10}\rho+\sqrt2h<\sqrt{3/2}-1.
$$

The stopped process therefore stays inside the rotating plateau. Transforming back bounds the endpoint well error by

$$
\rho+\sqrt{20/9}\,h.
$$

Equal-sign pairs remain in the inactive region on the ordinary Brownian good event, with endpoint error at most $\rho+h$.

Every active-unit bound is below $0.013$. During restoration, the noiseless path stays in its well. The stopped $2h$ comparison stays strictly inside the $0.1$ neighborhood where the restoring drift is nonincreasing, so there is no stopped exit. The restoration endpoint error is less than

$$
0.013e^{-8.55}+2h<0.004003<\rho.
$$

After each successful boundary, fresh conditional Brownian increments allow the same bound from every point in the new well box.

### Couple the complete finite episode

Union-bound the first failure over active and restoration intervals, all $m$ cells and all $G$ gates. The resulting episode bound is

$$
\delta_G\leq
\min\!\left\{
1,
G\left(
8m e^{-h^2/(2\sigma^2)}
+8e^{-h^2/(9\sigma^2)}
\right)
\right\}.
\tag{B.4}
$$

Preparation failure is an additional charge. On the common good event, the native decoded transcript equals the ideal transcript. Their total variation distance is therefore at most $\delta_G$. For two comparison arms, their success difference can change by at most the sum of their episode error bounds.

This coupling covers the specified powered bit primitives. A stochastic source ingress must have its exact declared law or a separate uniform conditional kernel-error bound. An internal-gate noise estimate does not certify an external sensor. Source errors, acquisition failures and clock failures must be included in the episode bound by the same first-failure argument before applying the inquiry comparison tolerance.

Equation (B.4) is a finite-horizon approximation. The unconditional sign process need not be exactly Markov: distinct voltages inside one well can have different later error probabilities. An exact SPC-2 claim for the stochastic voltage process requires a complete continuous/native closure analysis of that process. The error estimate alone cannot establish an exact finite quotient.

The distinction is useful in practice. A sufficiently small episode error can preserve the measurable inquiry advantage, while the stronger claim of exact native closure remains a separate physical obligation.

## B.4 The inquiry machine's full cost ledger

Use NAND decompositions with the following counts:

| Operation | NAND gates |
|---|---:|
| XOR | 4 |
| Multiplexer | 4 |
| AND | 2 |
| Three-bit majority | 6 |

The circuit stages have these charges:

| Stage | Computation gates |
|---|---:|
| Two initial endorsement copies | 2 |
| Agreement and initialization | 8 |
| Masked third endorsement read | 3 |
| Priority installation | 4 |
| Each flag-update stage | 17 |
| Final prediction | 14 |
| Requests for the $L-1$ later calibrations | $L-1$ |
| Request for the final target | 1 |

The total is

$$
2+8+3+4+17L+14+(L-1)+1=31+18L.
$$

There are $L+2$ inquiry slots of 27 ordinary windows and one final stage of 18 windows:

$$
27(L+2)+18=27L+72.
$$

Only two flags are needed because fewer than $L$ calibrations cannot reach the first useful threshold, and at exactly $L$ only all-positive or all-negative increment histories reach it. An eight-cell scratch allocation is initialized before use in every stage. The paper's circuit tables identify every source and target. These counts do not assert a minimum bank size.

With $m=23$, each identity tour costs

$$
2(m-1)=44.
$$

An initial tour and one ordinary-operation/tour pair per window give

$$
N_{\mathrm{native}}=44+45(27L+72)=1215L+3284.
$$

Each of the $L+1$ raw-triple sites has one joint native ingress boundary but costs three single-write work units. That difference adds $2(L+1)$ work units:

$$
N_{\mathrm{work}}=1215L+3284+2(L+1)=1217L+3286.
$$

An absent calibration writes constant zeros at its allocated site and retains the same cost allowance. Both comparison arms use the same padded resources. A saved inquiry opportunity can therefore be assigned to another task without giving the adaptive arm an uncharged physical implementation.

### Conditional time and energy bounds

Let $\tau_g$ bound an elementary gate window and $E_g$ bound a single-write work unit. Let $\tau_{\mathrm{port}}$ and $E_{\mathrm{port}}$ bound an acquisition interval and its energy charge. Then

$$
T\leq(1215L+3284)\tau_g+(L+3)\tau_{\mathrm{port}},
$$

$$
E\leq(1217L+3286)E_g+(L+3)E_{\mathrm{port}}+E_{\mathrm{aux}}.
$$

The auxiliary term includes charged clock, program, holding and other apparatus overhead. These are conditional resource upper bounds, not thermodynamic equalities.

A fixed 23-cell bank does not imply constant total apparatus storage. A literal unrolled program grows linearly with $L$, and its phase counter grows logarithmically. Every source and coupler state needed for joint ingress and subsequent laws must be modeled. A serial unobserved buffer cannot be left out of that inventory.

The result is an explicit route from an evaluative program to a costed physical episode. The mathematical mechanism, its information restrictions and its native returns can be preserved together. Assessing an actual device then requires checking the complete apparatus against this contract.


<!-- published: References; /consciousness/agency/references -->

# References and publication record

The numbered references below preserve the bibliography of **Bounded Agency and Reflective Freedom: Evaluative Revision, Information Limits, and Faithful Realization**. Citation numbers in its web chapters link here. Supplementary essays identify their own dependencies and do not assign the published paper credit for additional unpublished results.

The author-supplied publication is Jeremy Rodgers, Independent Researcher, 6 October 2026, Version 2.0. [DOI: 10.5281/zenodo.23202999](https://doi.org/10.5281/zenodo.23202999). [Read the supplied PDF](/publications/consciousness/agency/bounded-agency-and-reflective-freedom.pdf).

## 1. Abstracting causal models {#ref-1}

Sander Beckers and Joseph Y. Halpern. *Abstracting causal models.* Proceedings of the AAAI Conference on Artificial Intelligence, 33(1):2678–2685, 2019. [doi:10.1609/aaai.v33i01.33012678](https://doi.org/10.1609/aaai.v33i01.33012678). [Article](https://ojs.aaai.org/index.php/AAAI/article/view/4117).

## 2. Certified self-modifying code {#ref-2}

Hongxu Cai, Zhong Shao, and Alexander Vaynberg. *Certified self-modifying code.* Proceedings of the 28th ACM SIGPLAN Conference on Programming Language Design and Implementation, pp. 66–77. ACM, 2007. [doi:10.1145/1250734.1250743](https://doi.org/10.1145/1250734.1250743). [Author page and extended version](https://flint.cs.yale.edu/shao/papers/smc.html). Extended version: Yale Technical Report YALEU/DCS/TR-1379.

## 3. Games with imperfect information {#ref-3}

Krishnendu Chatterjee, Laurent Doyen, Thomas A. Henzinger, and Jean-François Raskin. *Algorithms for omega-regular games with imperfect information.* Logical Methods in Computer Science, 3(3), article 4, 2007. [doi:10.2168/LMCS-3(3:4)2007](https://doi.org/10.2168/LMCS-3(3:4)2007). [Article](https://lmcs.episciences.org/1094).

## 4. Self-modification of policy and utility {#ref-4}

Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter. *Self-modification of policy and utility function in rational agents.* In Bas Steunebrink, Pei Wang, and Ben Goertzel, editors, Artificial General Intelligence, Lecture Notes in Computer Science 9782, pp. 1–11. Springer, 2016. [doi:10.1007/978-3-319-41649-6_1](https://doi.org/10.1007/978-3-319-41649-6_1). [arXiv:1605.03142](https://arxiv.org/abs/1605.03142).

## 5. Responsibility and manipulation {#ref-5}

John Martin Fischer. *Responsibility and manipulation.* The Journal of Ethics, 8(2):145–177, 2004. [doi:10.1023/B:JOET.0000018773.97209.84](https://doi.org/10.1023/B:JOET.0000018773.97209.84). [Author PDF](https://andrewmbailey.com/jmf/Responsibility_and_Manipulation.pdf).

## 6. Responsibility and control {#ref-6}

John Martin Fischer and Mark Ravizza. *Responsibility and Control: A Theory of Moral Responsibility.* Cambridge University Press, 1998. [doi:10.1017/CBO9780511814594](https://doi.org/10.1017/CBO9780511814594). [Publisher](https://www.cambridge.org/core/books/responsibility-and-control/54D0EB8AEDEF4D5F4930D691EC214E01).

## 7. Sequential testing under deadlines {#ref-7}

Peter Frazier and Angela J. Yu. *Sequential hypothesis testing under stochastic deadlines.* Advances in Neural Information Processing Systems, volume 20, 2007. [Proceedings](https://proceedings.neurips.cc/paper/2007/hash/9c82c7143c102b71c593d98d96093fde-Abstract.html).

## 8. Adaptive submodularity {#ref-8}

Daniel Golovin and Andreas Krause. *Adaptive submodularity: Theory and applications in active learning and stochastic optimization.* Journal of Artificial Intelligence Research, 42:427–486, 2011. [doi:10.1613/jair.3278](https://doi.org/10.1613/jair.3278). [arXiv:1003.3967](https://arxiv.org/abs/1003.3967). Corrected author version: v5, 2017.

## 9. Transitive factorisations {#ref-9}

I. P. Goulden and D. M. Jackson. *Transitive factorisations into transpositions and holomorphic mappings on the sphere.* Proceedings of the American Mathematical Society, 125(1):51–60, 1997. [Author PDF](https://uwaterloo.ca/math/sites/default/files/uploads/documents/gjpams1997.pdf).

## 10. Smooth simulation of computation {#ref-10}

Daniel S. Graça, Manuel L. Campagnolo, and Jorge Buescu. *Robust simulations of Turing machines with analytic maps and flows.* New Computational Paradigms, Lecture Notes in Computer Science 3526, pp. 169–179. Springer, 2005. [doi:10.1007/11494645_21](https://doi.org/10.1007/11494645_21). [Author PDF](https://sqigmath.tecnico.ulisboa.pt/pub/GracaDS/05-GCB-stable.pdf).

## 11. Selecting computations {#ref-11}

Nicholas Hay, Stuart Russell, David Tolpin, and Solomon Eyal Shimony. *Selecting computations: Theory and applications.* Proceedings of the Twenty-Eighth Conference on Uncertainty in Artificial Intelligence, pp. 346–355, 2012. [Proceedings PDF](https://www.auai.org/uai2012/papers/123.pdf).

## 12. Adaptivity gaps in Boolean evaluation {#ref-12}

Lisa Hellerstein, Devorah Kletenik, Naifeng Liu, and R. Teal Witter. *Adaptivity gaps for the stochastic boolean function evaluation problem.* arXiv:2208.03810, version 1, 2022. [Preprint](https://arxiv.org/abs/2208.03810).

## 13. Discovering agents {#ref-13}

Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt. *Discovering agents.* Artificial Intelligence, 322:103963, 2023. [doi:10.1016/j.artint.2023.103963](https://doi.org/10.1016/j.artint.2023.103963). [Author PDF](https://sebastianfarquhar.com/assets/papers/kentonDiscovering2023.pdf).

## 14. Can AI systems have free will? {#ref-14}

Christian List. *Can AI systems have free will?* Synthese, 206:115, 2025. [doi:10.1007/s11229-025-05209-x](https://doi.org/10.1007/s11229-025-05209-x).

## 15. Unsupervised reliability calibration {#ref-15}

Yao Ma, Alex Olshevsky, Csaba Szepesvari, and Venkatesh Saligrama. *Gradient descent for sparse rank-one matrix completion for crowd-sourced aggregation of sparsely interacting workers.* Journal of Machine Learning Research, 21(133):1–36, 2020. [Article](https://jmlr.org/papers/v21/19-359.html).

## 16. Chance, choice and control {#ref-16}

Henry D. Potter and Kevin J. Mitchell. *Chance, choice, and control: free will in an indeterministic universe.* Synthese, 207:209, 2026. [doi:10.1007/s11229-026-05570-5](https://doi.org/10.1007/s11229-026-05570-5).

## 17. General agents need world models {#ref-17}

Jonathan Richens, Tom Everitt, and David Abel. *General agents need world models.* Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research 267, pp. 51659–51687. PMLR, 2025. [Proceedings](https://proceedings.mlr.press/v267/richens25a.html).

## 18. Relational boundaries and awareness localization {#ref-18}

Jeremy Rodgers. *Relational boundaries and awareness localization: Robustness, composition, and identification limits.* Zenodo, 2026. Preprint, version 1.0, 29 September 2026. [Full website treatment](/consciousness/research/paper-2).

## 19. Shadow Theory and Consciousness {#ref-19}

Jeremy Rodgers. *Shadow theory and consciousness: Awareness, perspectival realization, and the source-to-experience problem.* Zenodo, 2026. SPC-2, version 2, publication edition, 20 September 2026. [Full website treatment](/consciousness/monograph).

## 20. Learning effective interfaces {#ref-20}

Jeremy Rodgers. *Learning effective interfaces from opaque stochastic systems: Capacity, selection, and validation limits.* Zenodo, 2026. Revised preprint, version 2, 29 September 2026. [Full website treatment](/consciousness/research/paper-3).

## 21. Identifying binary realizations {#ref-21}

Jeremy Rodgers. *Identifying binary realizations from intervention laws: Certificates, recoding obstructions, and a bounded SPC-2/IIT comparison.* Zenodo, 2026. Preprint, version 1.1-RC1, 30 September 2026. [Full website treatment](/consciousness/research/paper-4).

## 22. Causal consistency {#ref-22}

Paul K. Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M. Mooij, Dominik Janzing, Moritz Grosse-Wentrup, and Bernhard Schölkopf. *Causal consistency of structural equation models.* Proceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, 2017. [arXiv:1707.00819](https://arxiv.org/abs/1707.00819).

## 23. Risk-aware partially observed planning {#ref-23}

Pedro Santana, Sylvie Thiébaux, and Brian Williams. *RAO\*: An algorithm for chance-constrained POMDP's.* Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pp. 3308–3314. AAAI Press, 2016. [doi:10.1609/aaai.v30i1.10423](https://doi.org/10.1609/aaai.v30i1.10423). [Article](https://ojs.aaai.org/index.php/AAAI/article/view/10423).

## 24. Gödel machines {#ref-24}

Jürgen Schmidhuber. *Gödel machines: Fully self-referential optimal universal self-improvers.* In Ben Goertzel and Cassio Pennachin, editors, Artificial General Intelligence, Cognitive Technologies, pp. 199–226. Springer, 2007. [doi:10.1007/978-3-540-68677-4_7](https://doi.org/10.1007/978-3-540-68677-4_7). [Author page](https://people.idsia.ch/~juergen/goedelmachine.html).


<!-- supplement: Extended inquiry 1; /consciousness/agency/bounded-perspective -->

# Bounded perspective, real control

An agent never needs to possess the whole world in order to act within it. What it needs depends on the task. A missing distinction matters when different possibilities demand incompatible responses. Other hidden distinctions can remain hidden without preventing success.

This is the first extension of the published account of reflective freedom. The published paper establishes how retained commitments can govern assessment, rule installation and fresh decisions. The supplementary control models developed here ask a different supporting question: **what can an embedded system accomplish through a limited interface?** Their results belong to this extended website treatment. They are not additional claims of the published paper.

## What a perspective preserves

Write the observational relation as

$$
\Omega \xrightarrow{\Delta} R.
$$

The source state is in $\Omega$, the aperture $\Delta$ supplies a readout, and $R$ is what the installed interface makes available. If $\Delta$ is many-to-one, several source states have the same readout. That is an exact limitation on discrimination. It does not yet specify a limitation on every possible action.

For a deterministic aperture, the three relevant tests are different:

| Question | Exact requirement |
| --- | --- |
| Can the source state be reconstructed? | $\Delta$ must be injective. |
| Can a particular fact $q(\omega)$ be recovered? | $q$ must be constant on each fibre of $\Delta$. |
| Can the task be completed without regret? | The source states in each fibre must share an optimal executable action. |

A fibre is the set of source states giving one readout. Knowing which fibre contains the world can be enough to answer one question and insufficient to answer another. A standpoint is therefore a structured access relation, not a scalar amount of ignorance.

Consider a target register holding an unknown value $x$ and a receiver prepared at a desired value $g$. The reversible swap

$$
(x,g)\longmapsto(g,x)
$$

sets the target exactly without identifying $x$. The unresolved value survives in the receiver. Conversely, even complete observation does not help a controller whose only permitted actuation is the identity. Adding an inaccessible spectator register changes the completeness of its world description while leaving this target task unchanged.

These examples block an unrestricted inference from partial knowledge to absent agency. They also identify the right questions: which actions are compatible with the remaining uncertainty, where can unresolved information go, and which transformations can the apparatus execute?

## The exact decision obstruction

Let $\Theta$ be a finite set of source conditions, $\mathcal H$ a finite record alphabet, and $Q(h\mid\theta)$ the observation channel. Let $\mathcal A$ be the admitted finite action set. A policy $\pi(a\mid h)$ can use the record but cannot read the hidden condition directly.

An outcome kernel $K(y\mid\theta,a)$ and a bounded task score $u(\theta,y)$ define

$$
v_\theta(a)=\sum_y K(y\mid\theta,a)u(\theta,y),
\qquad
v^*_\theta=\max_a v_\theta(a),
\qquad
\ell_{\theta a}=v^*_\theta-v_\theta(a).
$$

The regret $\ell_{\theta a}$ measures the task value lost by selecting $a$ when $\theta$ is the source condition. Every row has at least one zero. Define worst-source regret by

$$
R(Q,\ell)=\min_\pi\max_\theta
\sum_{h,a}Q(h\mid\theta)\pi(a\mid h)\ell_{\theta a}.
$$

**Theorem: exact finite task compatibility.** The same value has the dual expression

$$
R(Q,\ell)=\max_{\lambda\in\Delta(\Theta)}
\sum_h\min_a\sum_\theta\lambda_\theta Q(h\mid\theta)\ell_{\theta a}.
$$

Moreover, $R(Q,\ell)=0$ exactly when every possible record $h$ satisfies

$$
\bigcap_{\theta:\,Q(h\mid\theta)>0}
\operatorname*{arg\,min}_a\ell_{\theta a}\ne\varnothing.
$$

**Proof.** The policy set is a product of finite probability simplexes. Replace maximization over the source condition by maximization over source distributions $\lambda$. Finite minimax exchanges that maximization with policy minimization. With $\lambda$ fixed, the policy optimization separates over records, and a minimizing action at each record gives the dual formula.

If each intersection is nonempty, choose an action in it. Every supported source-record pair then has zero loss. Conversely, zero worst-source regret means that every nonnegative term with positive $Q(h\mid\theta)\pi(a\mid h)$ has zero loss. An action given positive probability at $h$ must be optimal in every source condition compatible with that record. At least one such action exists because the policy probabilities sum to one. This proves both directions.

The obstruction is **incompatible required action under the same available record**. Randomization does not remove that obstruction at zero error. Every action in a randomized policy's support must still satisfy the common requirement.

Pairwise agreement is weaker than joint agreement. With a single record and loss matrix $\ell=I_3$, any two source rows share a zero-loss action, but all three do not. The minimax value is $1/3$: assigning each action probability $1/3$ attains it, and some action must receive at least that probability. Testing only pairs would miss the obstruction.

## How indistinguishability becomes a quantitative limit

Total variation is

$$
\operatorname{TV}(P,Q)=\frac12\sum_h|P(h)-Q(h)|.
$$

It measures how well the available records can distinguish two hypotheses. Suppose two source conditions have disjoint optimal action sets and every nonoptimal action loses at least $\gamma>0$. Then

$$
R(Q,\ell)\ge
\frac\gamma2\left[1-\operatorname{TV}(Q_{\theta_0},Q_{\theta_1})\right].
$$

**Proof.** Their common record mass is $c(h)=\min\{Q(h\mid\theta_0),Q(h\mid\theta_1)\}$, with total mass $1-\operatorname{TV}(Q_{\theta_0},Q_{\theta_1})$. At any action, the two losses sum to at least $\gamma$ because the action cannot be optimal in both conditions. Sum over the common mass and take half the resulting total. Maximum risk is at least average risk.

For several source conditions, put

$$
\alpha=\sum_h\min_\theta Q(h\mid\theta),
\qquad
r_{\rm blind}=\min_p\max_\theta\sum_a p(a)\ell_{\theta a}.
$$

Then $R(Q,\ell)\ge\alpha r_{\rm blind}$. For $\alpha>0$, the common record mass induces the same action distribution in every condition; discard the nonnegative loss on the remaining mass. For $\alpha=0$, the statement is just nonnegativity.

Garbling an observation cannot improve the optimum when the actuator and resources stay fixed. If $Q'=QB$, every policy using $Q'$ can be simulated from $Q$ by first drawing the garbled record through $B$. Its policy class is contained in the original one. This finite argument is the relevant instance of [Blackwell's comparison of experiments](https://doi.org/10.1214/aoms/1177729032).

The lesson for reflective agency is direct. An acceptable amendment can exist without being identifiable through the installed read ports. The published paper's common-amendment criterion applies this same structure to endorsement and continued auditability. More information can remove a conflict between compatible contexts. It cannot supply an unavailable actuator or make an inaccessible amendment executable before a deadline.

## Capability is a set of achievable consequences

For a specified interface and policy-resource class $\Pi$, define

$$
\mathcal K(\Pi)=\{K^\pi(\cdot\mid\theta):\pi\in\Pi\}.
$$

This collects achievable outcome or transcript laws. It permits exact comparisons between a controller before and after learning, between two memory budgets, or between two operation sets. There is no need to force these comparisons into one universal number called freedom.

**Theorem: a task witnesses a capability gap.** Let $\mathcal C$ be a compact convex set of finite kernels, let $\mu$ have full support, and let $K$ be a target kernel. Then

$$
\begin{aligned}
\min_{L\in\mathcal C}\sum_\theta\mu_\theta\operatorname{TV}(K_\theta,L_\theta)
=\max_{0\le u\le1}\Bigg[&\sum_{\theta,y}\mu_\theta K(y\mid\theta)u(\theta,y)\\
&-\max_{L\in\mathcal C}\sum_{\theta,y}\mu_\theta L(y\mid\theta)u(\theta,y)\Bigg].
\end{aligned}
$$

**Proof.** For probability distributions of equal mass, $\operatorname{TV}(p,q)=\max_{0\le u\le1}\sum_yu_y(p_y-q_y)$. Apply this separately to each source row and exchange the minimum over the compact convex kernel class with the maximum over the score cube by finite minimax. The inner minimum subtracts the largest comparator score.

A new achievable law outside the old convex capability class therefore has a bounded task that reveals the difference. This is a precise sense in which learning can create an effective capability. It does not establish that the learner has escaped its underlying physical laws or acquired ultimate authorship of them.

## A world model can be partial and adequate

If estimated action values obey $|\widehat v_\theta(a)-v_\theta(a)|\le\varepsilon$ for every value actually compared, the greedy estimated choice loses at most $2\varepsilon$ relative to that comparison's true optimum. Add and subtract the two estimated values; the greedy inequality cancels the middle difference. The premise must cover the comparison being claimed. An error guarantee on accessible comparisons does not automatically reach an inaccessible source-oracle optimum.

Nor does a lossy observation automatically destroy Markov structure. For a controlled process on $X$ and a map $\phi:X\to Z$, an exact controlled quotient exists when

$$
\phi_*K^a(\cdot\mid x)=\phi_*K^a(\cdot\mid x')
\quad\text{whenever }\phi(x)=\phi(x'),\quad\text{for every }a.
$$

Necessity follows by starting from either state in the same fibre. Sufficiency defines the quotient kernel by the common pushed-forward law and then iterates it. Memory is required when relevant distinctions fail this test, not merely because the representation is incomplete.

Broad competence can demand much richer models than a single reset task. [General agents need world models](https://proceedings.mlr.press/v267/richens25a.html) supplies such necessity results under its own task and environment assumptions. A bounded perspective can support strong local competence without becoming a complete representation of the source. That is the setting in which reflective freedom has to be built and tested.


<!-- supplement: Extended inquiry 2; /consciousness/agency/reversible-storage-and-calibration -->

# Where uncertainty goes

Control changes a selected part of the world. In a reversible model, it must also preserve the distinctions carried by the initial state. A target can become orderly while information moves into a receiver, a retained record or another explicitly modeled output.

This supplementary analysis gives that statement an exact finite form. It extends the website's account of bounded agency with receiver and calibration results. The published reflective-freedom paper uses different devices and resource units; their counts should not be combined into a fictitious single machine.

## The receiver capacity theorem

Let a finite plant $x\in X$ have prior $\mu(x)$. A nondisturbing observation produces retained record $h$ through $Q(h\mid x)$. The receiver starts in one specified ready state and has $D$ possible final states. Conditional on $h$, the actuator can apply any permutation of plant and receiver. Additional scratch must return to fixed values; otherwise its possible final values count toward $D$.

Let the target set $G\subseteq X$ have $g$ elements. For a nonnegative weight vector $w$, write $\operatorname{Top}_s(w)$ for the sum of its $s$ largest entries, or all entries if there are fewer than $s$.

**Theorem: exact unrestricted receiver capacity.**

$$
p^*=\sum_h\operatorname{Top}_{gD}
\left((\mu(x)Q(h\mid x))_{x\in X}\right).
$$

For uniform input on $K$ states with a deterministic record partition $\{C_h\}$,

$$
p^*=\frac1K\sum_h\min\{|C_h|,gD\}.
$$

**Proof.** Fix $h$. Distinct ready-receiver inputs require distinct outputs under a permutation. Exactly $gD$ outputs place the plant in its target. Thus at most $gD$ inputs can succeed, and the largest possible successful mass is the sum of the $gD$ greatest weights. Conversely, map those inputs injectively into successful output positions and complete the partial injection to a permutation of the finite register space. Do this separately for each retained record. The record remains unchanged during actuation.

This is an exact count of available output positions. It is not a thermodynamic work law, and it does not promise an efficient circuit for every attaining permutation.

For $N$ independent uniform $K$-state inputs, a common $D$-state receiver and targets of size $g$, the blind bound is

$$
p_{\rm all}\le\min\{1,D(g/K)^N\}.
$$

Here **blind** means that no retained record carries source information. The joint receiver is the only output allowed to retain unresolved source distinctions. Success at least $1-\varepsilon$ consequently requires

$$
\log_2D\ge N\log_2(K/g)+\log_2(1-\varepsilon).
$$

An informative record changes the accounting. Record a fair bit exactly as $h=x$ and use the reversible CNOT $(x,h)\mapsto(x\oplus h,h)$. The plant resets with certainty and no extra residual receiver, because the record already contains the distinction. Applying the blind bound would be wrong. With a general retained channel, use the full conditional top-mass expression.

## The quantum capacity counterpart

For a density operator $\rho$ and a success projector $P$ of rank $s$,

$$
\max_U\operatorname{Tr}(PU\rho U^\dagger)=\sum_{j=1}^s\lambda_j^\downarrow(\rho).
$$

**Proof.** In an eigenbasis of $\rho$, put $w_j=\langle j|U^\dagger PU|j\rangle$. These weights satisfy $0\le w_j\le1$ and $\sum_jw_j=s$. The weighted eigenvalue sum is largest when the largest $s$ eigenvalues receive weight one. A unitary aligning their eigenvectors with the range of $P$ attains that choice.

A maximally mixed $K$-state input therefore has success at most $gD/K$ in a $gD$-dimensional success space. The statement uses the ordinary density-operator probability rule. It does not derive that rule or permit copying an arbitrary unknown quantum state. It identifies the same capacity obligation under unitary dynamics: unresolved distinctions require physical room.

## Learning an unknown sensor

Let $V=\mathbb F_2^d$, $K=2^d$, $d\ge2$, and suppose the unknown sensor is

$$
f(x)=Ax+b,\qquad A\in\operatorname{GL}(d,2),\quad b\in V,
$$

uniform over the affine group. A live input is uniform and independent of the sensor. Before it arrives, the apparatus can query $q$ chosen reference states and retain their readings. The task is exact regulation of the live plant to one prescribed value, with a $D$-state ready receiver and unrestricted record-conditioned permutations.

**Theorem: optimized calibration–receiver frontier.**

$$
p^*_{q,D}=\begin{cases}
\min\{1,D/K\},&q=0,\\
\min\{1,(2^{q-1}+D)/K\},&1\le q\le d,\\
1,&q\ge d+1.
\end{cases}
$$

This optimizes the reference design. For a particular nonempty transcript whose references have affine-span dimension $s$, the value is

$$
p^*_{s,D}=\frac{2^s+\min\{D,K-2^s\}}K.
$$

**Proof.** Without a reference, affine transitivity makes the source posterior uniform even after the live sensor reading. The capacity theorem gives $D/K$, capped at one.

After a reference, translate its source position to zero and subtract its observed offset. Let $W$ be the span of reference differences, with dimension $s$. The remaining sensor ambiguity is the group fixing $W$ pointwise. In a chosen complement its matrices have form

$$
\begin{pmatrix}I_s&B\\0&C\end{pmatrix},\qquad C\in\operatorname{GL}(d-s,2).
$$

Every point of $W$ is fixed. All points outside $W$ form one orbit: choose $C$ to map one nonzero complement coordinate to another, then choose $B$ to adjust the $W$ coordinate. A live source in $W$ is exactly known. Outside $W$, its posterior is uniform on $K-2^s$ points. The receiver preserves at most $D$ of those successful possibilities, giving the stated fraction.

Each new reference can add at most one independent difference. References $0,e_1,\ldots,e_{q-1}$ attain $s=q-1$ until $d+1$ references determine the sensor. Adaptive pre-live choice cannot exceed the same span bound. Conditional on a full calibration transcript, its choice rule adds no observation beyond the queried readings; the remaining uniform sensor posterior is a coset of the same pointwise stabilizer.

Repeated readings of the same reference do not count as new independent directions. At $d=2,D=1$, two references $(0,0)$ give success $1/2$, while distinct references $(0,e_1)$ give $3/4$. Counting operations without checking what they distinguish can overstate competence.

With $r$ receiver bits, $D=2^r$, this single-task optimum can be attained with conditional affine data operations. For $r\le d-1$, an appropriate $r$-flat fits in the unresolved complement and can be transferred to the receiver by affine normalization. A full $d$-bit receiver permits a blind swap. Shared tasks introduce more demanding geometry.

## Shared calibration creates correlated uncertainty

One sensor acting on $N$ live inputs leaves a common residual group $G$ acting diagonally on $V^N$. Conditional on the complete readings, the source tuple is uniform on a group orbit. Hence

$$
p^*_{N,G,D}=K^{-N}\sum_{O\in V^N/G}\min\{|O|,D\}.
$$

This is one shared receiver, not a fresh receiver for every task. After calibration spanning $s$ dimensions, put $m=d-s$. If the complement components of the source tuple have rank $k$, their orbit size is

$$
L_{s,k}=2^{sk}\prod_{j=0}^{k-1}(2^m-2^j).
$$

The product counts injective images of the $k$ independent complement directions. The factor $2^{sk}$ counts their possible $W$ components. The rank probability is

$$
\rho_{m,N}(k)=2^{-mN}\prod_{j=0}^{k-1}
\frac{(2^m-2^j)(2^N-2^j)}{2^k-2^j},
$$

with empty products equal to one. Counting rank-$k$ matrices by image subspace and full-rank coordinate map gives this expression. Applying the receiver bound orbit by orbit yields

$$
p^*_{N,s,D}=\sum_{k=0}^{\min(m,N)}\rho_{m,N}(k)
\min\{1,D/L_{s,k}\}.
$$

Without a reference, the corresponding affine-span orbit of $N$ points has size $2^d\prod_{j<k}(2^d-2^j)$, where $k$ is the rank of their $N-1$ differences.

## A deadline changes the task

A batch controller sees all readings before acting. A causal controller may have to release each plant before the next reading arrives. Equal total information does not guarantee equal available action at the earlier deadline.

There is an exact positive result under uniform symmetry. Let one uniform hidden $g\in G$ act on independent uniform plants, giving $Y_t=gX_t$. Retain every reading immutably. Only the current plant and one persistent $D$-state receiver may change during actuation; the sensor is isolated, no extra source-sensitive probe is allowed, and all additional helpers must be restored before release. The targets are singleton states.

**Theorem: online orbit capacity.** Under these conditions,

$$
p^*_{\rm online}=p^*_{\rm batch}=K^{-N}\sum_{O\in X^N/G}\min\{|O|,D\}.
$$

**Proof.** Batch capacity is an upper bound. For a fixed reading prefix, possible source prefixes form a uniform orbit. Projection onto the preceding prefix has equal-size fibres, because group elements biject the extensions of any two prefixes. If prefix-orbit sizes obey $L_t=L_{t-1}b_t$, keep up to $D$ successful prefixes in distinct receiver states. When $L_{t-1}\le D$, all prefixes survive, producing $L_t$ candidate extensions. When $L_{t-1}>D$, the retained $D$ prefixes have $Db_t\ge D$ extensions. In both cases select exactly $\min(D,L_t)$ extensions, map their distinct plant-receiver pairs to the target and distinct receiver states, and complete the map to a permutation. No released plant is touched again. Uniformity makes the retained fraction optimal at every final reading.

The receiver must remain available for interaction. A sealed archive has a different role. The proof gives exact existence; history-conditioned permutations may be expensive, and retaining every reading does not give bounded total memory.

For a general finite exogenous joint law $w(x_{1:N},y_{1:N})$, causal success instead has a nested-list characterization. At each reading history $h_t$, choose a set $L(h_t)$ of at most $D$ compatible source prefixes, beginning with the empty prefix, such that

$$
\{x_{1:t-1}:x_{1:t}\in L(h_t)\}\subseteq L(h_{t-1}).
$$

Maximize the terminal mass

$$
\sum_{h_N}\sum_{x_{1:N}\in L(h_N)}w(x_{1:N},h_N).
$$

**Why this is exact.** Successful source histories cannot merge into one receiver state when earlier plants and records are fixed. Thus every controller supplies such lists. Conversely, the distinct selected extensions can be assigned distinct receiver labels through partial permutations, completed at each step. Online and batch values agree precisely when a feasible list family captures the terminal top-$D$ mass at every positive-probability final reading. This characterization is finite, without an efficiency claim.

For a concrete gap, take $Y_i=X_i\oplus S$, with $S$ fair and independent $X_i\sim\operatorname{Bernoulli}(p)$, $p\ge1/2$. With $D=1$, the first causal reset commits to an interpretation and the optimal all-task success is $p$. Batch success is

$$
\frac12\sum_{y\in\{0,1\}^N}
\max\{p^{|y|}(1-p)^{N-|y|},p^{N-|y|}(1-p)^{|y|}\}.
$$

For $N=3,p=4/5$, the values are $4/5$ online and $112/125$ in batch. The improvement is purchased by waiting for later evidence. A capacity can exist in the architecture yet be unavailable on a particular occasion because its evidence arrives too late.


<!-- supplement: Extended inquiry 3; /consciousness/agency/transformations-and-control -->

# The operations that make control possible

An agent's information does not determine its competence by itself. The same records and the same amount of residual storage can support different achievements when the available transformations differ. Agency is embodied in an operation set as well as an observation channel.

The supplementary results here isolate that difference. Binary affine operations comprise invertible combinations of XOR, bit flips and swaps. A Toffoli adds a nonlinear product: it changes a target bit by the product of two control bits. Its role is computational. A Toffoli count is not a count of decisions, a measure of consciousness or an energy budget.

## The geometry of a successful set

Let the uncertain plant be $X\in\mathbb F_2^n$. After record $h$, write its unnormalized weight as $w_h(x)=\Pr(X=x,h)$. Let $r$ ready receiver bits remain variable on success, and let the permitted plant target be an affine set of dimension $t$. Every other helper and required output is fixed on success.

Define the largest posterior mass on a flat of dimension at most $b$ by

$$
\Phi_b(w)=\max_{\substack{L\subseteq\mathbb F_2^n\;\mathrm{affine}\\\dim L\le b}}
\sum_{x\in L}w(x).
$$

For $b\ge n$, this is the whole mass.

**Theorem: exact affine optimum.** With record-conditioned affine reversible data operations,

$$
p^*_{\rm aff}=\sum_h\Phi_{r+t}(w_h).
$$

**Proof.** An affine bijection maps the ready-input plane into another affine plane. Intersecting it with the success constraints has an affine preimage. Only $r+t$ output directions can vary on success, so that preimage has dimension at most $r+t$. Conversely, take any permitted flat of that dimension, normalize it by invertible affine coordinates, transfer its varying directions to the target and receiver, and complete to an affine bijection. The maximizing flat attains the bound.

Unrestricted reversible actuation can select the $2^{r+t}$ most probable individual inputs. Affine actuation must fit its successful inputs inside an affine flat. Their arrangement matters even when their number is small enough.

## The restricted Clifford counterpart

The same optimum holds for unitary Clifford operations when the input is a computational-basis ensemble, the ready auxiliaries are independent pure stabilizer states, and the target/receiver test is the corresponding stabilizer projector:

$$
p^*_{\rm Cliff}=p^*_{\rm aff}.
$$

**Proof.** Conjugating the success projector by a Clifford gives a stabilizer projector. In its Pauli expansion, only diagonal Pauli terms contribute to a computational-basis diagonal entry. Their binary constraints define either the empty set or an affine set $L$. The acceptance profile is therefore $\alpha\mathbf1_L$ for $0<\alpha\le1$.

Restricting to the prepared input plane gives trace at most the success projector's rank, so $\alpha|L|\le2^{r+t}$. If $\dim L\le r+t$, its weighted acceptance is at most the mass of an admitted flat. Otherwise partition $L$ into parallel $(r+t)$-flats. The largest part has at least the average mass, which bounds $\alpha w_h(L)$. Affine circuits are Clifford circuits and attain the classical optimum.

The stabilizer structure is established in [Dehaene and De Moor](https://doi.org/10.1103/PhysRevA.68.042318). This application excludes extra measurements, nonstabilizer preparations, uncounted residual outputs and arbitrary coherent source ensembles. Its conclusion is an operation-class comparison under that precise contract.

## A six-point separation

Take two two-bit live plants behind one unknown affine sensor, with one calibration reference. Removing the observed offset leaves a common $\operatorname{GL}(2,2)$ ambiguity. The source-pair space has one rank-zero state, three rank-one orbits of size three, and one full-rank orbit of size six.

In a four-bit encoding the full-rank orbit is

$$
S=\{6,7,9,11,13,14\}.
$$

Its largest intersections with affine flats of dimensions $0,1,2,3,4$ have sizes

$$
1,\quad2,\quad3,\quad4,\quad6.
$$

Why can a two-dimensional plane contain only three? Four plane points XOR to zero. All six listed words XOR to zero, so a zero-XOR subset of four would leave two distinct words with zero XOR, which is impossible. Any three distinct binary points span a plane and attain three. A hyperplane contains at most four of these six; a hyperplane containing five would contain the sixth because their total XOR is zero, but the affine hull is four-dimensional. Suitable hyperplanes attain four.

With $r=2$ receiver bits, affine and Clifford control reach

$$
\frac{1+3\min(2^r,3)+\max_{\dim L\le r}|L\cap S|}{16}
=\frac{13}{16}.
$$

Unrestricted control reaches

$$
\frac{1+3\cdot3+4}{16}=\frac78.
$$

One nonlinear data operation closes the gap. Translate by $9$ and apply the invertible linear map with columns $(8,14,1,5)$. Four selected states become $0,1,2,7$. A Toffoli with controls $0,1$ and target $2$ changes $7$ to $3$ while fixing the other three. They now lie on a coordinate plane that can be moved into the receiver.

This is one extra nonlinear operation on uncertain data after the record is fixed. It is not a claim that every other part of the sensing apparatus is affine. The distinction prevents a local operation-class result from being misreported as a comparison of two entirely different kinds of computer.

## An exact family of storage–computation tradeoffs

For $m\ge1$, let the plant be uniform on the graph

$$
\Gamma_m=\{(u,v,q_m(u,v)):u,v\in\mathbb F_2^m\},
\qquad q_m(u,v)=\sum_{i=1}^m u_iv_i.
$$

There are $2^{2m}$ source points in $2m+1$ plant bits. The controller must reset every plant bit to a prescribed target. It has $r$ ready receiver bits, arbitrary affine reversible operations, at most $k$ ordinary classical Toffolis, and extra helpers that are restored on success.

**Lemma: flat intersection.** Every affine $l$-flat in $\mathbb F_2^{2m+1}$ meets $\Gamma_m$ in at most

$$
\min\{2^l,2^{l-1}+2^{m-1}\}
$$

points.

**Proof.** Project onto $(u,v)$. If projection is not injective on the flat, every projected point has two vertical lifts and exactly one is on the graph, giving $2^{l-1}$. Otherwise the graph condition on the projected flat is a quadratic equation plus an affine term. Its polar form is the restriction of

$$
\beta((u,v),(u',v'))=u\cdot v'+u'\cdot v.
$$

On an $l$-dimensional direction space, this restriction has rank at least $2l-2m$: its radical lies in an orthogonal complement of dimension $2m-l$. Splitting a binary quadratic character sum into hyperbolic pairs shows that its magnitude is zero or $2^{l-R/2}$ for polar rank $R$. It is therefore at most $2^m$. The number of zeros is at most half the flat size plus half this magnitude, giving $2^{l-1}+2^{m-1}$. The cardinality bound supplies the other term.

**Theorem: exact quadratic-source frontier.** For $0\le r\le2m$ and $k'=\min(k,m)$,

$$
p^*_{m,r,k}=2^{-2m}\min\{2^r,\;2^{r-1}+2^{m+k'-1}\}.
$$

**Upper bound.** Run a successful circuit backward from its $r$-dimensional output plane. At each inverse Toffoli, split one control into the cases zero and one. On either branch the gate is affine. After at most $k$ gates, successful inputs lie in at most $2^k$ disjoint affine pieces $L_i$, whose total size is at most $2^r$. The lemma bounds the graph points by

$$
\frac12\sum_i|L_i|+2^k2^{m-1}
\le2^{r-1}+2^{m+k-1}.
$$

Injectivity independently bounds successful inputs by $2^r$. For $k\ge m$, that cardinality bound is sufficient.

**Attainment.** Cancel $k'$ disjoint products from the graph bit with $k'$ Toffolis. Then transfer $r$ base coordinates into the receiver, constraining the rest to zero. Choose the $2k'$ coordinates of cancelled products first, then one member of each remaining pair, then the other members. If $r\le m+k'$, no surviving product has both coordinates free and all $2^r$ assignments succeed. Otherwise exactly $j=r-m-k'$ uncancelled product pairs remain free. Their inner product has $2^{2j-1}+2^{j-1}$ zero assignments. Multiplication by the other free-bit factor gives $2^{r-1}+2^{m+k'-1}$. The circuit uses $k'$ Toffolis and $3r$ CNOTs for swaps, with no extra scratch.

At the tight receiver width $r=2m$,

$$
p^*_{m,2m,k}=\frac12+2^{k-m-1},\qquad0\le k\le m.
$$

Exactly $m$ Toffolis are necessary and sufficient for perfect control at that width. With $m-1$, the exact optimum is $3/4$. One additional receiver bit allows a wholly affine swap of the complete plant and gives perfect control.

This is a scalable substitution between residual storage and nonlinear processing. The source preparation itself is a resource. Arbitrary coherent Hadamard interleavings are outside the positive-Toffoli proof because classical affine branching no longer describes their amplitudes.

## A reusable envelope bound

The same argument applies whenever a source distribution obeys, for every affine flat $L$,

$$
\sum_{x\in L}p(x)\le a|L|+b,\qquad a,b\ge0.
$$

With an exact target, $r$ receiver bits, restored workspace and $k$ Toffolis,

$$
p_{\rm success}\le\min\{1,\operatorname{Top}_{2^r}(p),a2^r+b2^k\}.
$$

To prove it, use the same inverse affine partition. Intersect each piece with the prepared input plane, project to the source coordinates and apply its mass bound. The piece sizes sum to at most $2^r$, and at most $2^k$ additive terms $b$ appear. Injectivity supplies the independent top-mass bound.

These formulas show why practical freedom has several axes. Better evidence, more writable storage and richer transformations can each alter what is achievable. None can be inferred solely from the presence of the others. The published evaluator adds the further requirement that the available transformations are selected through the declared evaluative process and actually installed for later use.


<!-- supplement: Extended inquiry 4; /consciousness/agency/constructed-interfaces -->

# Building a new way to act

A controller can learn how its sensor works, construct a decoder and install a new way of acting. After that installation, it can handle inputs it could not previously manage. The change is real even though the complete learning process follows fixed physical rules.

The published paper makes this point for evaluative rules: an installed charter can become an object of assessment and be changed for fresh later cases. This supplementary construction concerns a different object, a learned sensor interface. It provides a concrete account of capability acquisition and exposes the separate costs of building, storing and using that capability.

## Acquiring a source-preimage basis

Let a fixed unknown invertible binary sensor supply

$$
Y_t=AX_t,\qquad A\in\operatorname{GL}(d,2).
$$

Each live plant gives one stipulated nondisturbing reading. Its value must then be regulated to a prescribed target and released before the next reading. Sensor parameters are isolated. Reading records are retained immutably and cannot become an extra actuation target. Additional source-sensitive probes are forbidden. All temporary helpers must be restored at release.

The controller maintains an echelon basis of the public readings. For each basis vector $E_j$, its physical receiver holds the corresponding preimage

$$
R_j=A^{-1}E_j.
$$

When a reading $y$ arrives, reduce it against the stored reading basis. Apply the same XOR coefficients from receiver blocks to the current plant. If the reading was dependent, the plant becomes zero. If it adds an independent direction, swap the remaining physical plant into a ready receiver block and store the corresponding reading residual in the public basis. Finally XOR the desired goal into the plant and release it.

The invariant $R_j=A^{-1}E_j$ proves the procedure. It does not require the controller to receive $A$ as hidden advice. Physical source values enter the receiver through the admitted plant operations, while public readings organize how those values are subsequently used.

A specified affine streaming compiler uses $d^2$ receiver bits, $3d^2+6d+1$ working bits apart from receiver and archive, and

$$
N(12d^2+21d)
$$

elementary gates for $N$ tasks. Per task, its counts are $2d$ NOTs, $6d^2+13d$ CNOTs and $6d^2+6d$ Toffolis; the retained reading/dispatch archive has $4Nd$ bits. These are compiler bounds, not optima. The Toffolis include sensor and record-controlled processing even when the uncertain-data operation at a fixed record is affine. A nonzero affine offset needs separately charged calibration.

## The exact price of one saved receiver bit

After $d$ independent observed directions, the uncertain source-preimage basis is a matrix $B\in\operatorname{GL}(d,2)$. Its query convention is $Bc$, where $c$ gives coefficients in the public reading basis. It is not an automatically supplied oracle for the hidden sensor or its inverse under another convention.

**Theorem: exact residual receiver width.** For $d\ge2$, a horizon admitting $d$ independent directions and the one-readout, isolated-sensor, immutable-archive contract above, perfect regulation for every admitted sensor and source sequence requires and permits

$$
r_{\rm affine/Clifford}=d^2,
\qquad
r_{\rm unrestricted}=d^2-1.
$$

**Proof of the lower bounds.** The number of possible source-preimage bases is

$$
|\operatorname{GL}(d,2)|=\prod_{j=0}^{d-1}(2^d-2^j)
=2^{d^2}\rho_d,
\qquad
\rho_d=\prod_{j=1}^d(1-2^{-j}).
$$

For $d\ge2$, $1/4<\rho_d\le3/8<1/2$, so

$$
\left\lceil\log_2|\operatorname{GL}(d,2)|\right\rceil=d^2-1.
$$

The lower constant can be seen without a numerical infinite product: retain the first three factors and bound the remaining product below by one minus the sum of its omitted $2^{-j}$ terms. It gives a bound exceeding $9/32$. Distinct source bases must remain distinct when the plants have all reached fixed targets, giving the unrestricted bit requirement.

The affine hull of $\operatorname{GL}(d,2)$ is the whole $d^2$-dimensional matrix cube. For every off-diagonal entry, $I$ and $I+E_{ij}$ are invertible and differ only there. For a diagonal entry, use a swapped invertible two-by-two block; toggling that entry preserves invertibility. These differences span every coordinate direction. An affine success preimage containing all invertible matrices must therefore have dimension $d^2$. The affine/Clifford flat theorem gives the stronger width bound.

**Attainment.** The streaming basis construction uses $d^2$ bits. For unrestricted actuation, the clean encodings below fit the full basis in $d^2-1$ bits. Before the last independent reading, at most $d(d-1)$ preimage bits are stored, which fits this receiver. The final live plant supplies the last column temporarily; encode the completed basis before releasing it. Later actions use its compact representation.

The saved bit concerns residual source-dependent storage at release. It does not include all temporary workspace, public basis records, source parameters, clock or program memory.

## A clean compact code

A clean one-bit encoder is a fixed NOT/CNOT/Toffoli circuit on $d^2$ raw matrix bits and ready auxiliaries. On every invertible input, exactly $d^2-1$ designated outputs may vary; all others have prescribed constants. The complete circuit remains a permutation on every input, including singular matrices. The successful cleanup promise applies to invertible matrices unless a stronger domain is stated.

One recursively usable code takes $B=[b\mid Z]\in\operatorname{GL}(j,2)$. Let $P_b$ swap the first nonzero coordinate of $b$ into position zero. With $b'=P_bb$, define

$$
(L_{b'}z)_0=z_0,
\qquad
(L_{b'}z)_i=z_i+b'_iz_0\quad(i>0).
$$

Then $F_b=L_{b'}P_b$ sends $b$ to $e_0$, giving

$$
F_bB=\begin{pmatrix}1&a\\0&C\end{pmatrix},
\qquad C\in\operatorname{GL}(j-1,2).
$$

Define

$$
E_j(B)=(b,a,E_{j-1}(C)).
$$

Its length is $j+(j-1)+((j-1)^2-1)=j^2-1$. At $j=2$, using column-major bits $0,1,2,3$, the gate sequence

$$
T(0,3;1),\quad CX(1;3),\quad T(2,3;1),\quad CX(1;3),\quad X(3)
$$

clears the fourth output on all six invertible inputs. The remaining three bits form the base code. Here $T(a,b;t)$ adds the product of control bits $a,b$ into target $t$, $CX(c;t)$ adds control $c$, and $X(t)$ flips $t$.

The decoder reconstructs $C$, forms the displayed block matrix and applies $F_b^{-1}$. To make the encoder physically clean, evaluate $E$ while keeping $B$, copy its code, and uncompute the evaluator's intermediates. Then use the retained code to XOR the decoded $B$ out of the original input positions. This yields $(0,E(B),0)$ on the promise. It does not attempt to uncompute an operation after destroying the input its inverse needs.

This is the established reversible compute/copy/uncompute method associated with [Bennett's logical reversibility](https://doi.org/10.1147/rd.176.0525), applied to the specified code and cleanup contract.

## Installation and use are different problems

Assume uniform exact binary matrix multiplication has Boolean circuits of size $O(n^\beta)$ for a fixed $\beta>2$. The recursive code can be cleanly installed with

$$
O(d^\beta+d^2\log^3d)
$$

Toffolis and clean temporary bits. The admitted seven-product matrix recursion gives the explicit conservative choice $\beta=\log_2 7$.

The construction block-factorizes the first $d-2$ columns while preserving chronological first-nonzero pivots. At each split it factors the left block, permutes the right block, solves through the unit-lower factor and forms a Schur residual by multiplication. Later row swaps are reversed on the retained multiplier histories to recover the recursive code; they do not require repeating the whole factorization for each column. Fixed sorting networks charge the routing overhead at $O(d^2\log^3d)$. Decoding reconstructs the factors with the same order. Clean Boolean evaluation supplies the reversible implementation.

The same code supports

$$
(E(B),c,x,0)\longmapsto(E(B),c,x+Bc,0)
$$

with $O(d^2)$ Toffolis and $O(d)$ query helpers. For $c=(c_0,v)$,

$$
Bc=F_b^{-1}\begin{pmatrix}c_0+av\\Cv\end{pmatrix}.
$$

Compute the smaller product once, continue upward, copy the completed result into $x$, and reverse the entire computation. The sum of level costs is quadratic. Computing and uncomputing the smaller query twice at each recursive level would introduce an unjustified larger cost. One explicit compiler for this code uses $9d^2-13d+6$ Toffolis and $3d-1$ clean helpers; smaller constants from a different representation cannot be silently attached to its fast installation bound.

**Theorem: programmable query lower bound.** Any fixed classical circuit computing $Bc$ for every represented invertible $B$ and every runtime $c$ needs

$$
k\ge\frac{d(d-1)}2
$$

ordinary Toffolis, regardless of the chosen classical representation or clean auxiliary width. Thus the best exact programmable-query order is $\Theta(d^2)$.

**Proof.** Let Alice know $B$ and its representation, and Bob know $c$. Represent every circuit wire by XOR shares. Affine gates are local. For a Toffoli, Alice sends her two control shares; Bob can use these together with his own shares to update the product share correctly. With $k$ Toffolis this uses $2k$ bits, all from Alice. Alice sends her $d$ final output shares, letting Bob recover $Bc$.

Alice's message depends on $B$, not $c$. If two different operators produced the same message, Bob would obtain the same product for every basis input $c$, forcing the operators to agree. Hence $2k+d\ge\log_2|\operatorname{GL}(d,2)|$. Integer rounding gives the displayed bound. The recursive query supplies the matching upper order.

This concerns one programmable circuit handling every coefficient word. A specialized fixed query, such as $c=0$, is another task. For any clean encoder with cost $k$, decoding, raw matrix-vector accumulation and re-encoding also give the useful bound

$$
Q_E(d)\le2k+d^2.
$$

## Why even one-bit compression has a quadratic nonlinear cost

Let $C_{\rm enc}(d)$ be the minimum Toffoli count of any clean encoder under the contract above. Then

$$
C_{\rm enc}(d)=\Omega(d^2).
$$

The lower bound uses a named external theorem: in the additive two-party input model over a fixed prime field, distinguishing rank $d$ from rank $d-1$ with public-coin error at most $1/10$ needs $\Omega(d^2\log p)$ communicated bits. This is the rank theorem of [Li, Sun, Wang and Woodruff](https://arxiv.org/abs/1407.4755). Its deep proof is imported; the reduction to the encoder follows here.

Let $Z$ be the set of raw matrices whose constrained encoder outputs all take the required constants. Invertible matrices belong to $Z$, while injectivity gives $|Z|\le2^{d^2-1}$. The density of rank-$d-1$ matrices is

$$
q_d=(2-2^{1-d})\rho_d.
$$

Consequently

$$
\Pr(B\in Z\mid\operatorname{rank}B=d-1)
\le\frac{1/2-\rho_d}{q_d}<\frac{14}{27}.
$$

The constraints accept every full-rank matrix and reject at least $13/27$ of the adjacent rank class. Public uniform invertible left and right multipliers make an XOR-shared input uniform within its rank class, and each party transforms its own share locally. Simulate the encoder with two communicated bits per Toffoli. A public random parity of the constrained-output deviations detects any nonzero deviation with probability $1/2$, using one additional parity-share bit. One repetition has miss probability at most $41/54$ on rank $d-1$. Nine independent repetitions reduce it below $1/10$ and communicate at most $18k+9$ bits. The imported rank theorem forces $k=\Omega(d^2)$.

The asymptotic constant is unspecified. This is not the statement $C_{\rm enc}(d)\ge d^2$ with coefficient one in every small dimension. Together with the construction, it gives

$$
\Omega(d^2)\le C_{\rm enc}(d)\le O(d^\beta+d^2\log^3d).
$$

The missing general near-quadratic installation theorem remains open. A controller can build a compact, reusable capability; knowing that such a capability exists does not make its construction free.


<!-- supplement: Extended inquiry 5; /consciousness/agency/perspective-history-and-sourcehood -->

# A chooser with a history

A decision can be the agent's decision even though the agent has a history. The substantive question is how that history acts now: which commitments remain operative, which can be examined, what evidence is available, and whether an assessment can change the rule used next.

The published paper answers part of this question constructively. Retained commitments are read, candidate rules are assessed, a selected rule is installed and fresh cases reveal the changed disposition. The philosophical development here connects that result to the supplementary mathematics of provenance, capability construction and bounded perspective. Its interpretations extend the discussion; they do not add unproved conclusions to the published theorems.

## The chooser as the operating perspective

The agent's perspective is where available evidence, retained commitments and executable possibilities meet. The proposal is to locate choosing in that organized process. A further entity watching the process would need its own means of receiving information and making changes. Adding it would postpone the explanatory work.

This account still has to distinguish an evaluator from a switch. A command changing when an internal bit changes demonstrates dependence. It does not establish that the command agrees with the relevant commitments, that the bit was obtained through a legitimate read, or that a revised disposition was installed. The published inverted-selector example shows why: a faithful selector and a systematically reversed one can have identical causal contrasts.

The physical mechanism, the evaluative relation and the interpretation of a perspective therefore have different jobs. The mechanism supplies causal operations. The declared relation says which amendments accord with the retained commitments. The perspective interpretation concerns the organized agent to which those operations are attributed. None of them can be replaced by simply naming a register "self" or "value."

For a person, the relevant perspective includes an embodied history, learned expectations, attention and practical circumstances. The finite constructions isolate some of those functional relationships. They do not prove the presence of phenomenal awareness in a bit circuit. Under the website's SPC-2 account, experiential assignment requires the separate native constitution and its correspondence laws.

## Conditioning can remain as content without ruling every act

An old policy can remain recorded while losing its control over present action. This is a stronger and more useful distinction than treating freedom as the destruction of one's history.

Let $S$ encode an old policy and $L$ a newly installed one. On a ready archive register, the mapping

$$
(S,L,0)\longmapsto(L,L,S)
$$

is injective and can be extended to a reversible permutation of the full register space. Subsequent actions based on the first register depend on $L$, and, with $L$ fixed, no longer depend on $S$. The old policy survives in the archive. A later authorized audit might consult it, but its mere existence no longer gives it compulsory authority over current deliberation.

This construction does not prove that $L$ was appropriately chosen. For that, attach the published assessment, endorsement and installation tests. It shows that release from an old influence and preservation of its record are physically compatible.

There is an exact information identity behind this separation. Suppose that, conditional on public context $E$, the final controller $C$ and residual system $B$ retain enough information to reconstruct an earlier source variable $S$. Then

$$
H(S\mid E)=I(S;C\mid E)+I(S;B\mid C,E).
$$

**Proof.** Reconstruction gives $H(S\mid C,B,E)=0$. Apply the chain rule to $I(S;C,B\mid E)$, which therefore equals $H(S\mid E)$.

If the active controller contains at most $\varepsilon$ bits about $S$, at least $H(S\mid E)-\varepsilon$ bits remain conditionally in the residual system. Historical information can move outside the active decision path without disappearing from the complete physical state. Statistical information, present causal influence and ultimate origination are separate properties.

This is also the right distinction for discussing the constructed self. Personal memories, social roles and practical self-descriptions can remain useful content. The philosophical question is whether they must serve as the unquestioned center from which every interpretation and response is organized. The RCO hypothesis, where RCO means recursively closed observer, addresses a reorganization of that authority. The bounded-agency constructions show particular forms of assessable rule revision; they do not establish the full RCO hypothesis or identify its transition in a human subject.

## What the present can reveal about the past

Two different histories may leave the same accessible boundary state while differing in an inaccessible archive. If the boundary gives the observer the same future interaction laws, further questioning cannot extract a difference that those laws do not carry.

**Theorem: accessible-boundary provenance bound.** Let two histories induce boundary distributions $\mu_0,\mu_1$ and the same subsequent controlled transition kernels. For every fixed adaptive audit policy $\pi$,

$$
\operatorname{TV}(P_0^\pi,P_1^\pi)
\le\operatorname{TV}(\mu_0,\mu_1).
$$

With equal prior probabilities, every binary provenance test has error at least

$$
\frac{1-\operatorname{TV}(\mu_0,\mu_1)}2.
$$

**Proof.** Compose the policy and common controlled kernels into a stochastic channel from boundary state to complete transcript. Total variation contracts through a common channel. For equal binary priors, the optimal discrimination error is one half minus one half the transcript total variation. Combining the two statements proves the bound.

Identical accessible laws are thus indistinguishable to every such audit. This is compatible with globally reversible history. Distinct complete pasts need not merge into one complete present state; their differences can remain outside the nominated boundary. Granting access to the archive changes the premise and can reveal provenance.

The published paper states a corresponding limit for identical complete present-state distributions and future boundary laws. It also develops the positive response: record the acquisition of commitments across a declared developmental interval. Observe the proposal, the information inspected, the assessment, the actual installation and its effect on fresh cases. Exclude or identify bypass writes. This provides evidence of audit-mediated acquisition during that interval.

Even a faithfully operating assessor can be misled by a controlled evidential environment. Suppressing contrary observations or fabricating apparently credible evidence can preserve every internal audit pathway. Current responsiveness therefore does not, by itself, distinguish transparent persuasion from deceptive information supply. Provenance has to be investigated at the sources and channels where that difference exists.

## A new capability need not break a law

Calibrating an unknown affine sensor gives a concrete example. References at $0,e_1,\ldots,e_d$ identify its offset and columns. A learner can retain those records, compute a decoder and install it. The active controller now has an effective repertoire it previously lacked.

An equally equipped universal controller, given the same observations, storage, time and permission to run that procedure, can also execute it. Expanding the installed repertoire does not exceed a comparator class that already includes its construction. Conversely, a preinstalled correct decoder contains sensor-correlated information; calling it a cost-free baseline would hide the acquisition resource.

Both statements can hold together. **Learning changes what this controller can presently do. Fixed laws can govern that change.** A person learning a skill need not rewrite physics for the skill to become a new practical possibility.

The distinction has a formal normal form.

**Proposition: effective construction on a fixed carrier.** If states, instruments and programs have finite descriptions, and transition and installation rules are computable, the whole process can be represented on one fixed countable state space of finite words.

**Proof.** Choose effective encodings with tags for program, data, apparatus and records. Concatenate the current descriptions into one finite word. A universal interpreter implements each allowed transition, including writing a new program or apparatus description. The set of all finite words is fixed even when the process reaches descriptions absent from its former active repertoire.

The result concerns computably described processes. It does not declare every physical or experiential possibility computable. Within its scope, effective novelty is compatible with a fixed mathematical carrier and fixed transition law.

## Lawful does not mean predictably exhaustible

There is a sharp difference between checking a finite budget and deciding every future capability of a universal system.

Take a fixed universal controller with a protected latch. The latch reveals a fair hidden bit only after a supplied program halts. The admitted primitives prevent directly opening it. Before opening, best correct bit control is $1/2$; afterward it is one. Deciding whether any finite amount of internal computation eventually enables an improvement of at least $1/4$ would decide whether the supplied program halts.

**Proof of the obstruction.** Transform any program into the stated controller instance. A total procedure deciding eventual capability gain would answer its halting question. The diagonal halting argument rules out such a total procedure: a program could use its predicted halting behavior to do the opposite on its own description.

A fixed finite horizon with an explicitly finite action alphabet and finite kernels can instead be exhaustively evaluated. The unbounded theorem does not make every instance difficult, and it does not prevent discovering a particular new skill. It establishes that lawful self-development need not admit one universal complete forecast of every eventual capability.

## Ultimate authorship is a further demand

One proposed standard of ultimate authorship requires every qualifying act of self-constitution to be authorized by a strictly earlier qualifying act. Under a well-founded earlier-than relation, that requirement contradicts the existence of any qualifying act.

**Proof.** A nonempty set of qualifying acts has a minimal element. The requirement gives that element a strictly earlier qualifying authorizing act, contradicting minimality.

The inference is exact; the philosophical premise is disputed. Partial observation does not imply that every act must have this sort of prior authorization. The result cannot be substituted for a physics theorem excluding every account of agency.

The published paper deliberately permits fixed metarules. It asks whether an agent can assess and revise an operative rule, not whether it created every condition by which the assessment has meaning. That is a substantive local capacity with possible limits, rather than an answer obtained by demanding an impossible infinite ancestry of choices.

Randomness does not settle the issue either. Every finite stochastic kernel can be represented as a deterministic function of its input and an independent uniform seed by partitioning the unit interval. Rational probabilities permit finite noise tapes. This operational representation does not supply a local hidden-variable model for every quantum experiment, and it does not show that biological choices are deterministic. It shows why a stochastic transition table alone cannot decide ultimate authorship.

Even complete action-conditioned marginals need not fix individual counterfactuals. The models $(Y_0,Y_1)=(U,U)$ and $(U,1-U)$, for fair $U$, agree on the distribution observed under either action but disagree on whether the unchosen outcome equals the chosen one. Claims about that cross-action relation require an additional coupling assumption.

## Freedom across occasions

Freedom in this treatment has an operative profile: the process through which evaluation mediates choice, the consequence-distinct actions recruitable now, and the rules open to assessment and revision under the available budget. Information, integration and reasons responsiveness can guide philosophical evaluation of that profile. The finite mathematics does not supply a canonical moral score for them.

A commitment can protect a valued future and close another route. A self-binding decision can exhibit reflection now while reducing revisability later. Retaining a rule after examination can express responsiveness as clearly as changing it. A conflict can remain real after a decision has been settled. These cases require the time, target and evaluative relation to be specified rather than assuming that more options or more frequent change always means more freedom.

The resulting position gives self-governance something concrete to do. An agent can investigate its situation, assess inherited rules, redirect attention, install a new disposition and act through it. Its history constrains that process without making every historical influence permanently authoritative. Questions of moral responsibility, phenomenal identity and ultimate sourcehood remain additional arguments, with their own evidence and premises.


<!-- supplement: Extended inquiry 6; /consciousness/agency/paid-search-and-physical-realization -->

# The work behind an available option

An option is available to a bounded agent when its installed organization can identify and execute it in time. The existence of a mathematical answer is only the beginning. Finding the control instruction, routing its data, clearing temporary work and maintaining a physical record all have costs.

The supplementary compact-interface programme gives this distinction an unusually sharp form: an encoder can change at most one input bit at its endpoint while the computation selecting that bit remains a substantial problem. The same lesson applies to reflective inquiry. A simple final amendment need not be cheap to justify or implement.

## One changed bit can hide a difficult decision

Write an invertible binary matrix as $B=[b\mid H]$, where $H$ has $d-1$ columns and full column rank. Its unique nonzero left normal is $n(H)$, satisfying

$$
n(H)^TH=0,
\qquad
[b\mid H]\in\operatorname{GL}(d,2)\iff n(H)\cdot b=1.
$$

Fix $n\ne0$ and let $q=n\cdot b$. If $n_0=1$, define

$$
y_0=1+q,\qquad y_j=b_j\quad(j>0).
$$

If $n_0=0$, let $p>0$ be its first supported coordinate and set

$$
y_0=1+q,\qquad y_p=b_p+b_0+q,
\qquad y_j=b_j\quad(j\notin\{0,p\}).
$$

Keep $H$ unchanged and call this map $P(B)$.

**Theorem: direct half-cube encoding.** With a deterministic nonzero normal convention on deficient $H$ as well, $P$ is a permutation of the entire matrix cube. It sends every invertible matrix into the half-cube whose $(0,0)$ entry is zero, and changes at most one raw bit on every input. Omitting the zero coordinate leaves $d^2-1$ code bits.

**Proof.** For each fixed $H$, the normal is fixed. If $n_0=1$, the coefficient of $b_0$ in $y_0$ is one, making the first expression a bijection changing only that coordinate. If $n_0=0$, $q$ is independent of $b_0$ and has coefficient one on $b_p$. The two effective coordinates transform as

$$
(b_0,q)\longmapsto(1+q,b_0).
$$

This is bijective. When $b_0=q$, only coordinate zero changes; otherwise only coordinate $p$ changes. On the valid half $q=1$, the distinguished output is zero. The inverse for $n_0=0$ is recovered by $b_0=n\cdot y$, $q=1+y_0$ and $b_p=y_p+b_0+q$. The other case solves its affine first-coordinate equation. Thus every fibre is permuted and the whole map is reversible.

A total normal can be defined by binary elimination on a copy of $H$, applying the same invertible row operations to $I_d$. At each column choose the first remaining nonzero pivot, or its scheduled diagonal row when none exists, then clear below it. The last row of the transformed $H$ is zero, while the corresponding row of the transformed identity is nonzero. It supplies the normal even on deficient inputs.

Both $P$ and its inverse move every point by at most one. Therefore distinct inputs obey

$$
d_H(Px,Px')\le d_H(x,x')+2\le3d_H(x,x').
$$

Stored codewords differing in $t$ bits decode to matrices differing in at most $t+2$ entries; their actions on one coefficient word differ in at most $t+2$ output positions. These are endpoint fault-propagation bounds, not error correction or protection against faults during the calculation.

The number of possible rewrite locations still matters. If an encoder leaves every raw position outside a fixed set $R$ unchanged and removes one degree of freedom inside $R$, then $|R|\ge d$. To see this, any fewer than $d$ selected matrix entries can be left free while fixing the others so every filling is invertible. Choose a row with no selected entry, give it a unique one in a column containing a selected entry, delete that row and column and recurse. Determinant expansion proves the construction. If $|R|<d$, the resulting $2^{|R|}$ valid fillings would have to fit into $2^{|R|-1}$ output patterns, contradicting injectivity.

One input-dependent change and $d$ possible change locations are compatible. Neither fact says that the location is easy to determine.

## A supported coordinate supplies a control instruction

Let $A(d)$ be the minimum Toffoli cost of a deterministic clean selector returning a one-hot $e_j$ with $n_j(H)=1$ for every full-column-rank $H$. The selector preserves its input and restores additional helpers. It may choose any supported coordinate, without having to recover the entire normal or its first support position.

Given a selected $j$ and a first column $b$, form

$$
G(H,b,j)=\begin{pmatrix}H&0&b\\0&H&b+e_j\end{pmatrix}.
$$

If $m=n\cdot b$, its unique nonzero normal is

$$
n(G)=((1+m)n,mn).
$$

**Proof.** The first $2d-2$ columns span $\operatorname{im}H\oplus\operatorname{im}H$. The last column's two membership parities differ by $n_j=1$, so it lies outside that span. A normal has form $(\alpha n,\gamma n)$ and must satisfy $\alpha m+\gamma(m+1)=0$. Its nonzero solution is $(\alpha,\gamma)=(1+m,m)$. The normal's weight is unchanged.

Every supported coordinate of this larger matrix lies in the first block when $m=0$ and in the second when $m=1$. A support selector therefore also computes membership, after paying for the lifted matrix and cleanup. Compute the small selector, prepare the lift by affine copies, compute the larger selector, copy its block parity and reverse both evaluations. If $K_M$ is the clean membership-evaluator cost,

$$
K_M(d)\le2A(d)+2A(2d).
$$

Using that membership bit to route the selected first-column coordinate into the omitted slot gives a clean fibre encoder with

$$
C_{\rm enc}(d)\le2A(d)+4A(2d)+2d-1.
$$

The extra terms pay for controlled swaps and cleanup. The reduction needs an always-correct selector on every promised lifted input. A routine that can abstain is not interchangeable with it.

If the normal has weight one or two, the problem has a near-quadratic solution. Weight one means a zero row; weight two means two equal rows. A fixed sorting network on rows with original indices finds either witness in $O(d^2\log^2d)$ gates. Compute, copy the selected answer and uncompute. A unique three-row dependence instead asks for a distinct triple with zero XOR. Pairwise distinctness alone no longer identifies it.

## Identification, action and positive verification have different costs

Suppose an initial observation leaves a finite family $\mathcal F$ of nonempty possible supports. The only extra oracle answers whether a proposed support equals the actual one. Put

$$
M=|\mathcal F|,
\qquad
\Delta=\max_i|\{T\in\mathcal F:i\in T\}|.
$$

**Theorem: exact candidate-equality query costs.** The deterministic worst-case counts are

| Required result | Queries |
| --- | --- |
| Return any supported coordinate | $M-\Delta$ |
| Identify the whole support | $M-1$ |
| Obtain an actual positive oracle response | $M$ |

**Proof.** On an all-negative branch, each useful query excludes at most one possible support. Before $M-\Delta$ exclusions, more than $\Delta$ remain, so they cannot share a coordinate; no coordinate is safe to return. Conversely, choose a coordinate contained in $\Delta$ supports and query every support avoiding it. A positive answer provides a coordinate. If every answer is negative, the chosen coordinate must be supported. This uses $M-\Delta$ queries. Identification needs all but one possibility excluded and may infer the last; positive confirmation requires checking that final possibility too.

For all triples on $n$ unread indices, $M=\binom n3$, $\Delta=\binom{n-1}2$, and supported-coordinate selection needs $\binom{n-1}3$ equality queries. With four unread indices, one query of the triple avoiding a fixed index suffices: a positive response gives an answer; a negative response certifies the fixed index. Demanding positive verification would charge four.

These counts apply to the stated equality oracle. Reading a full nonzero residual or constructing a new observation from the raw matrix gives another information interface. A lower bound must name which interface and which output it concerns.

## A sound certificate is not a completed search

For a full-rank $d\times(d-1)$ matrix with a unique weight-three normal, let raw rows be $r_i$ and short hashes $h_i=\Lambda r_i$. The true triple always has hash XOR zero. False triples may do so too. Checking a candidate against its three full rows makes a successful answer sound. Failure to find one can mean abstention rather than absence of the true triple.

A monotone verifier chooses, at each anchor $a$, the lexicographically first distinct partner pair $(j,k)$ satisfying $h_a+h_j+h_k=0$, then checks all selected full residuals. Candidate order uses original indices. Adding parity rows removes false candidates without changing the order of those remaining, so an accepted true candidate stays accepted under refinement.

For a true anchor, acceptance occurs exactly when all lexicographically earlier false partners have nonzero projected residual. A random independent $\ell$-row parity map therefore gives, at any fixed true anchor,

$$
\Pr(\text{accept})\ge1-\left(\binom{d-1}{2}-1\right)2^{-\ell}.
$$

This follows by a union bound over nonzero blocker residuals. With $\ell=O(\log d)$, short sorting, paid full-row gathering and compute/copy/uncompute give a sound near-quadratic clean certificate. An expected number of retries does not prove a deterministic finite worst-case selector. Seeds and failed-trial histories also have to be retained or uncomputed.

A deterministic row-and-zero-injective sketch can be constructed at $O(d^2\log^3d)$ cost. Include one virtual zero alongside the distinct nonzero raw rows, so separating all supplied values makes every row hash nonzero as well as distinct. Choose each parity coordinate from low raw bit to high, assigning every unresolved pair to its highest differing raw coordinate and choosing the coefficient that leaves fewer such pairs unresolved, with zero on ties. Each parity halves the collision count. Grouped sorting and scans implement the pair counts without materializing all full-width pair differences. Yet an injective sketch only distinguishes rows. It need not distinguish a false triple from the true triple.

For example, raw rows

$$
(1,2,4,8,12,16)
$$

have true support $\{2,3,4\}$. The specified injector produces distinct nonzero hashes $(1,2,3,4,7,6)$, but earlier false partner pairs block every true anchor. The verifier abstains. More accurate observations must address the relevant relations, not just individual identities.

## Progress must be measured at the level a proof needs

The refinement statements below start from a row-and-zero-injective sketch and retain that property by appending parity rows. One refinement separates the currently selected nonzero full residuals from zero. A stronger finite-list refinement makes distinct selected residuals and zero mutually distinct. Neither condition says that the new sketch is injective on their whole linear span.

If a failed finite-list batch has $s$ active rows and $m$ distinct selected triples, then $m\ge\lceil s/3\rceil$. The distinct residuals all lie in the old sketch kernel. Giving them and zero distinct new images adds at least $\lceil\log_2(m+1)\rceil$ rank. Every deleting refinement also permanently removes the least active row: every projected triple through it is first at its partners, and all those selected false triples are eliminated on failure.

Charging these rank increases gives a universal $O(d/\log(d+1))$ failed-batch bound for the stated finite-list refinement. This ensures eventual exactness after a dimension-bounded unrolling. It does not establish near-quadratic total work as the growing sketch must be recomputed, routed and eventually cleaned.

A constructed family makes this gap decisive. For every $q=2^p\ge4$, the specified deterministic initializer and low-to-high, zero-tie residual refinement have an input with

$$
d_q=q^3+\binom q2,
\qquad
F(d_q)=q-3=\Theta(d_q^{1/3})
$$

failed batches. Its core rows encode graph edges with binary field labels $(i+j,i^3+j^3,i^5+j^5)$; projected zero-XOR triples correspond to triangles. Ordered coordinate blocks force the refinement to remove one graph vertex's star per stage, and guard rows force the actual initializer to preserve that pattern. The unique true dependence is the final triangle. This is a negative result for that precisely specified pipeline. The same inputs have a readily describable support and do not establish hardness for every selector.

Whole-span refinement gives a different guarantee with a different cost. Let $V_t$ be the span of all surviving projected-triple residuals, with dimension $v_t$, and let $R_t$ be the span of selected residuals. If every failed update is injective on $R_t$, then

$$
v_{t+1}\le\frac45v_t.
$$

**Proof.** Order distinct triples globally lexicographically. Every triple selected as the first at some vertex introduces that vertex relative to all earlier triples; hence the distinct selected incidence vectors are independent. The raw row map has one-dimensional kernel, so $m$ selected triples have residual rank at least $m-1$. They cover all $s$ active vertices, giving $m\ge\lceil s/3\rceil$, while $v_t\le s-1$. On failure, the three true anchors select distinct false triples, so $\dim R_t\ge2$. Thus

$$
\dim R_t\ge\max\{2,\lceil(v_t+1)/3\rceil-1\}\ge v_t/5.
$$

An update injective on $R_t$ removes at least that many dimensions from $V_t$'s surviving kernel. Therefore $v_{t+1}\le v_t-\dim R_t\le4v_t/5$.

There are $O(\log d)$ failed rounds, but the selected span can have $\Theta(d)$ dimensions. Computing and retaining a map injective on it has a larger implementation obligation. The safe clean bound is $O(d^3\operatorname{polylog}d)$, not a near-quadratic theorem.

A capped verifier instead lists projected triples using short pair hashes, orders them colexicographically by $(k,j,i)$ and checks only its first $K$ candidates against full rows. With $K=d$ and logarithmic hash width, its charged gate order is $O(d^2\log^3d)$. It immediately repairs the constructed slow family because its fewer than $q^3\le d$ core triangles occur before guarded candidates. It can still miss another input's true triple and must report abstention or overflow. An exact deterministic near-quadratic selector for every unique three-row dependence remains unresolved.

The practical implication is broader than this particular search. An agent can receive more information, make provable progress and still lack a sufficient budget for a completed decision. Round count, transcript width, routing cost and cleanup cannot stand in for one another.

## From a logical controller to a physical experiment

An ideal reversible gate list can be given finite Hamiltonians. For a Hermitian involution $U$ and gate duration $\tau$,

$$
H_U=\frac{\pi\hbar}{2\tau}(I-U)
\quad\Longrightarrow\quad
e^{-iH_U\tau/\hbar}=U.
$$

The eigenvalues of $U$ are $\pm1$; the two phases are respectively $1$ and $-1$, proving the identity. External sequencing remains a control resource.

A finite clock can implement a circuit autonomously at a chosen endpoint. For gates $U_0,\ldots,U_{L-1}$ and clock states $|0\rangle,\ldots,|L\rangle$, set

$$
H=\frac{\pi\hbar}{2T}\sum_{j=0}^{L-1}\sqrt{(j+1)(L-j)}
\left(|j+1\rangle\langle j|\otimes U_j+|j\rangle\langle j+1|\otimes U_j^\dagger\right).
$$

Conjugating by the block-diagonal history unitary removes the gate labels and leaves an engineered spin-chain clock. Its endpoint transfer at time $T$ applies the complete circuit, up to global phase. This uses the established [perfect-state-transfer construction](https://doi.org/10.1103/PhysRevLett.92.187902).

An endpoint is not a permanent record. The finite clock can evolve backward. Identity padding yields a binomial clock-position law with parameter $p(t)=\sin^2(\pi t/(2T))$ and can make completed computation likely at specified later observation times, but it does not automatically satisfy streaming release deadlines or preserve an uninterrupted classical transcript. The Hamiltonian norm is not a measured heat or work budget.

The published paper supplies a separate, more specific native realization: protected-source returns, successor-revealing records and executive access restrictions for its evaluators. The reversible-control clock here is supporting implementation mathematics, not a replacement SPC-2 certificate.

## Transfer the whole experiment

**Theorem: uniform instrument-level transfer.** Suppose the implemented initial state differs from its ideal state by at most $\eta_0$ in half trace norm. At stage $j$, suppose every history-conditioned implemented instrument differs from its ideal instrument by at most $\delta_j$ in half diamond norm, uniformly over the admitted controller class, including correlated memory and auxiliary references. Then

$$
\operatorname{TV}(P_{\rm impl},P_{\rm ideal})\le\eta_0+\sum_j\delta_j.
$$

The same bound controls the change in expected value of every complete-path score in $[0,1]$. For a candidate and baseline with respective path-error bounds $\eta_A,\eta_B$,

$$
g_{\rm impl}\ge g_{\rm ideal}-\eta_A-\eta_B.
$$

**Proof.** Retain the complete classical history and quantum or classical memory. Replace the initial state and then one instrument at a time in a hybrid sequence. Subsequent completely positive trace-preserving evolution contracts trace distance. Each instrument replacement contributes at most its diamond bound, even with a correlated reference. The triangle inequality sums the terms. Reading the final transcript is another channel and cannot increase the distance. A bounded path score changes by at most transcript total variation. Apply the bound separately to the candidate and baseline scores.

For a fixed pulse, Duhamel's formula gives the useful sufficient estimate

$$
\|U-\widetilde U\|\le\hbar^{-1}\int\|H(t)-\widetilde H(t)\|\,dt.
$$

It does not include preparation failure, clock correlations, leaked records or an unmodeled sensor. An optimized baseline comparison also needs a uniform error guarantee over every allowed baseline, rather than a check on one convenient circuit. The complete experiment must preserve the source law, observation restrictions, intervention routes and record lifetime for the logical advantage to become a physical claim.

The agent's effective freedom is consequently supported by a chain of explicit capacities: information it can actually obtain, transformations it can actually construct, amendments it can actually install, and consequences it can actually sustain. The published account supplies evaluative tests for that chain. These supplementary results show why its resource and realization conditions carry real explanatory weight.
