0 cumulative citations
View corpus contextA decision-theory for probing learners shows that deeper, compute-matched 'productive' probes can reveal action-changing hidden learning state and improve utility; tests on two 7B Transformer families confirm positive decision value and reusable compute advantages, though the empirical effect is system- and budget-specific.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.
Summary
Main Finding
Revelation Control formalizes when and how priced future-learning interventions (probes) should be used to reveal hidden, decision-relevant learner state — explicitly separating pure information value from the value of computation that the probe itself produces and can be reused. The paper gives exact decision-theoretic criteria (including a binary aliasing identity and a cost-adjusted scalar-control factorization), a concrete equal-budget geometry for comparing “restart” vs “promote” probe technologies, a sharp certification boundary for expected-utility claims, and cross-model empirical evidence (Qwen and Mistral 7B families) that deeper, productive probes can strictly improve decision value under equal compute.
Key Points
- Revelation Control problem: choose which priced probe to run (depth, protocol) so that any revealed distinctions are only as fine as needed to change the consequential decision; account separately for any useful progress left behind by the probe.
- Decision-sufficient revelation: refined information changes Bayes decision value only when it separates states inside a coarse-information fiber that favor different terminal actions.
- Binary aliasing identity (binary menu): refinement value = E[min{a_H, b_H}] — strict value exists exactly when a coarse-information fiber retains refined posterior mass on both sides of the terminal decision boundary.
- Revelation depth: the minimal probe depth at which a decision-relevant latent direction becomes first-order visible to the probe. Local geometry (derivatives of frontier map, gap, and probe readout) characterizes when a deeper probe reveals directions blind to a shallower probe.
- Productive vs restart probes:
- Restart (temporary probing): run short probe(s), then restart selected action from original anchor to full horizon (probe work discarded).
- Promote (productive revelation): continue the selected tested trajectory from probe endpoint — probe work is reusable.
- Equal-budget identity: for m candidate actions, an equal-budget productive horizon h matches a restart probe depth c when h = (m/(m-1)) c (for m=2, h = 2c). This quantifies the compute tradeoff between probing depth and reuse.
- Dynamic embedding and control:
- Static Bayes refinement value equals the population average of the state-dependent continuation (the value of buying extra depth conditional on current info).
- Conditional stopping rule: continue to deeper probe h' iff conditional expected gain G_{h→h'} exceeds priced compute cost c_{h→h'}.
- Cost-adjusted scalar-control factorization: a declared scalar shallow summary S_h suffices for the Stop/Continue decision exactly when, conditional on S_h, no scalar fiber contains states on both sides of the priced stop/continue boundary. The exact loss from using S_h rather than full shallow info equals E[min{a(S_h), b(S_h)}].
- Certification and learnability limits:
- Separates oracle-value from approximation/estimation/promotion terms.
- Shows bounded “stop–flip” risk alone cannot certify positive expected utility under unrestricted severity of outcomes; a severity/moment/tail condition (integrability) is necessary for robust population-level utility guarantees.
- Empirical instantiation:
- Two 7B Transformer families (Qwen2.5–7B and Mistral-7B-v0.3) were tested under an H12 terminal horizon, comparing restart H4 probes vs productive H8 probes with equal policy-visible update budgets (20 updates).
- Findings replicated across families: deeper probes can provide positive decision value; productive promotion yields strict equal-compute utility advantages.
- Qwen shows evidence of a decision-nonredundant shallow revealability regime (additional shallow signals matter). Mistral’s results are consistent with scalar decision sufficiency at the tested architecture family/resolution: a scalar continuation architecture fit on a development panel retained positive lower bounds on an independent target panel.
- Scope/limitations:
- Results are structural: decision theory, cost accounting, continuation logic, and evaluation protocol transport across models; fitted proxies, thresholds, coefficients, and required shallow-state dimensions may be model- and regime-specific.
- Finite-library and regime assumptions matter; population utility certification requires tail/severity assumptions.
Data & Methods
- Theory:
- Formal probability-space decision formulation; Bayes observation-relative value V_B(I) = E[max_a E[Q_a | I]].
- Definitions and propositions: decision factorization, dynamic obstruction decomposition, binary aliasing identity, revelation depth via differential geometry, equal-budget productive/restart identity, embedding of static Bayes refinement into state-conditioned continuation value, and cost-adjusted scalar-control factorization theorem.
- Proofs and statistical inference relegated to appendices.
- Empirical protocol (model-specific instantiation):
- Model families: Qwen2.5–7B and Mistral-7B-v0.3 (Transformer 7B-class).
- Controlled probe depths mapped to a common terminal horizon (H12); restart probes at H4 vs productive promote probes at H8 under equal total updates (20 policy-visible updates).
- Two-bank panel design: independent development and target panels to separate fitting of continuation architectures from evaluation; familywise-adjusted lower bounds used for statistical claims.
- Metrics: estimated Bayes decision value differences, decomposition into restart-quality and productive-reuse components, quantification of revealability regime (whether scalar summaries suffice), and robustness checks across independent panels.
- Empirical claims deliberately confined to these Transformer families and the tested horizons; implementation details and inference procedures provided in appendices.
Implications for AI Economics
- Value of experimental compute must separate information value from productive-computation value:
- When probes are promotable (their computation can be reused), their ROI can be substantially higher than restart probes for the same visible compute budget — this affects how practitioners and platforms should price and allocate exploratory compute.
- Equal-budget geometry gives a principled way to compare different probing protocols:
- Use h = (m/(m-1)) c to set depth tradeoffs that make restart and promote protocols directly comparable under equal visible compute; economists can use this to compare opportunity cost across experimental designs.
- Optimal stopping/control should be priced and implemented conditional on state, not by population quantiles:
- The conditional continuation value (G_{h→h'}) and the priced cost define the exact optimal continuation rule. Scalar summaries can be sufficient, but only when no scalar fiber straddles the priced Stop/Continue boundary — otherwise richer state monitoring is economically valuable.
- Risk and certification matter for expected-utility claims:
- Simple bounded-risk diagnostics (e.g., bounded stop–flip risk) are insufficient for guaranteeing positive expected utility under heavy tails; economic analyses and regulators should require explicit tail/severity/moment conditions when certifying policies that buy extra compute/experimentation.
- Practical deployment and governance:
- For model development, hyperparameter search, and model selection, accounting for promotion (continuing promising runs) vs restart changes the optimal allocation. Platforms and teams should explicitly account for whether probe computation will be reused when budgeting and pricing experiments.
- For high-stakes or safety-sensitive settings, revelation-control logic highlights a dual concern: probes both reveal hidden state and change the system. Governance should evaluate both informational benefits and the externalities/risks from promoting changed states into deployment.
- Transferability and empirical practice:
- Structural tools (decision-theory, cost accounting, continuation logic, evaluation protocol) can be adopted across model families. However, empirical proxies and thresholds (e.g., what scalar suffices, significance cutoffs) must be validated per-model and per-regime — economics analyses should allow for model-specific calibration.
- Recommended empirical-economic checklist for experiment design:
- Explicitly classify probe technology: restart vs promote. Account for reuse value when pricing.
- Match compute budgets via the equal-budget relation to compare protocols fairly.
- Compute conditional continuation values and compare to priced costs; implement state-dependent stopping.
- Test whether low-dimensional (scalar) summaries suffice for the meta-stop decision on independent development data before deploying cheaper control schemes.
- Require tail/severity checks (or moment assumptions) before claiming positive expected utility across populations.
Summary takeaway: Revelation Control provides a concise, operational decision theory and cost-accounting framework for when and how to run priced learning probes — crucial for economically optimal experiment/compute allocation — and shows both theoretically and empirically that productive (promoteable) deeper probes can strictly improve decision outcomes under equal compute budgets, subject to model- and regime-specific revealability and tail conditions.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| For finite action sets and integrable utilities, the Bayes value of a refined information set equals the value of the coarser information set if and only if a refined-information-optimal decision can be implemented using only the coarser information. Decision Quality | null_result | Whether additional information changes the optimal action and Bayes decision value |
Reading fidelity
high
Study strength
high
|
not reported
|
| For a binary action menu, additional information has strictly positive Bayes refinement value exactly when the refined posterior action-gap distribution retains positive mass on both sides of the terminal decision boundary within coarse-information cells. Decision Quality | positive | Bayes value gained from revealing hidden decision-relevant state |
Reading fidelity
high
Study strength
high
|
not reported
|
| A deeper future-learning probe can reveal a decision-changing hidden direction that a shallower probe cannot reveal, provided the shallow observation is locally insensitive to that direction while the deeper observation is locally sensitive. Decision Quality | positive | Decision-relevant revelation of hidden learning-state distinctions |
Reading fidelity
high
Study strength
high
|
not reported
|
| Under the paper's two-action, 12-horizon compute contract, a temporary H4 probe and a productive H8 probe consume the same 20 policy-visible updates. Organizational Efficiency | null_result | Policy-visible compute consumed under matched probing strategies |
Reading fidelity
high
Study strength
high
|
20 policy-visible updates for each method
|
| The value difference between productive evaluation of an active policy and restart evaluation of a frontier policy can be exactly decomposed into a restart-policy-quality term plus a productive-path-reuse term. Organizational Efficiency | mixed | End-to-end utility difference between productive and restart revelation technologies |
Reading fidelity
high
Study strength
high
|
not reported
|
| The population-average value of adaptively purchasing a deeper revelation is exactly equal to the static Bayes refinement value between the shallow and deeper information filtrations. Decision Quality | positive | Expected value of deeper information for subsequent decisions |
Reading fidelity
high
Study strength
high
|
not reported
|
| For a fixed shallow-to-deep policy pair, continuing the probe is pointwise optimal exactly when the conditional expected decision gain exceeds the additional compute cost. Task Allocation | positive | Net utility from choosing continuation rather than stopping |
Reading fidelity
high
Study strength
high
|
not reported
|
| A scalar shallow summary is decision-sufficient for the Stop/Continue meta-action if and only if no positive-probability scalar fiber contains states on both sides of the priced continuation boundary. Task Allocation | null_result | Loss in optimal stopping decisions from compressing shallow information to a scalar |
Reading fidelity
high
Study strength
high
|
not reported
|
| Bounded stop–flip risk alone cannot certify positive expected utility under unrestricted severity; an additional severity, moment, or integrable-tail condition is required. Decision Quality | negative | Ability to certify positive expected utility from stopping-risk bounds |
Reading fidelity
high
Study strength
high
|
not reported
|
| Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes showed positive decision value, and productive reuse produced strict equal-compute utility advantages. Organizational Efficiency | positive | Decision value of deeper probes and utility under productive reuse at equal compute |
Reading fidelity
high
Study strength
medium
|
positive decision value; strict equal-compute utility advantages
|
| In Qwen2.5-7B, the paper reports evidence that additional shallow revealability information is decision-nonredundant. Decision Quality | positive | Incremental decision value of additional shallow revealability information |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In Mistral-7B-v0.3, a scalar continuation architecture fit on an independent development panel retained positive familywise-adjusted lower bounds on a disjoint target panel. Decision Quality | positive | Out-of-sample lower-bounded utility of scalar continuation control |
Reading fidelity
high
Study strength
medium
|
positive familywise-adjusted lower bounds
|