The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Two different notions of informational superiority under sequential, evidence-dependent sampling coincide: one experiment can reproduce any stopping-based experiment and match decision value across finite decision problems if and only if it has larger directed KL divergences in both state directions; moreover, strict KL dominance guarantees higher value once observation costs are sufficiently small.

The Order of Binary Experiments under Endogenous Stopping
Zihao Li · August 20, 2026
arxiv theoretical high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Zihao Li unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Zihao Li provider ID
For finite binary experiments with endogenous stopping, the ability of one source to simulate another and its uniform value in costly-observation decision problems are equivalent and characterized exactly by coordinatewise dominance of the two directed Kullback–Leibler divergences, with constructive simulators attaining the KL bounds up to a source-dependent additive constant.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We study the comparison of binary statistical experiments in large samples where information is acquired sequentially and the number of observations can depend on realized evidence. We introduce two orders. Stopping dominance compares experiments by their ability to reproduce the outcomes of arbitrary stopping policies, while decision dominance compares their value in finite decision problems with costly observations. Our main result shows that the two orders coincide and are characterized by coordinatewise dominance of the two directed Kullback--Leibler divergences. Moreover, strict dominance implies eventual strict value dominance in every decision problem in which learning the state can change the optimal action. The key result behind this characterization is an exact simulation theorem: any binary target experiment can be generated from repeated observations of a source experiment with expected sample sizes attaining the two KL lower bounds up to an additive constant that depends only on the source.

Summary

Main Finding

The paper characterizes when one binary information source (experiment) is strictly better than another for sequential, evidence-dependent sampling. It introduces two orders — stopping dominance (ability to reproduce any stopping-generated experiment) and decision dominance (higher value in all finite decision problems with costly observations) — and proves they coincide. Both are exactly characterized by coordinatewise dominance of the two directed Kullback–Leibler (KL) divergences: F ⪰s G ⇔ F ⪰d G ⇔ I0(F) ≥ I0(G) and I1(F) ≥ I1(G), where Iθ(·) = KL(·θ ∥ ·1−θ). Moreover, strict dominance (both inequalities strict) implies eventual strict value dominance in any decision problem where learning can change the optimal action (for sufficiently small per-observation cost).

A core technical result is an exact simulation theorem: any finite binary target experiment H can be simulated from repeated (sequential, stop-rule) observations of source F with statewise expected sample sizes that meet the KL lower bounds Iθ(H) ≤ Iθ(F) Eθ[N] and, conversely, there exists a single simulator achieving the bounds up to an additive constant depending only on F (log of the ratio of extreme likelihood-ratios).

Key Points

  • Two comparison orders:
    • Stopping dominance (F can reproduce any experiment generated by stopping on G with at most statewise additive increase in expected sample size).
    • Decision dominance (F attains at least G’s value in every finite sequential decision problem with per-signal cost c, up to an O(c) additive term).
  • Equivalence and characterization: Both orders are equivalent and fully characterized by two directed KL inequalities I0 and I1.
  • Exact simulation bounds:
    • Any simulation must satisfy Iθ(H) ≤ Iθ(F) Eθ[N] (information lower bound).
    • There exists a simulator with Iθ(F) Eθ[N] ≤ Iθ(H) + constant(F) simultaneously for θ = 0,1. The constant is log(r̄F / r_F) where r̄F and r_F are the extreme pointwise likelihood-ratio values of the source.
  • Small-cost asymptotics: In sequential decision problems with small per-observation cost c, the optimal expected sample size scales like log(1/c)/Iθ(F), and value differences are O(c log(1/c)). Thus the KL coordinates govern asymptotic behavior.
  • Comparison to fixed-sample-size regime: Allowing endogenous stopping reduces the relevant information constraints to only the two directed KLs; by contrast, fixed (equal) sample-size comparison uses the full Rényi divergence spectrum (Mu et al., 2021). So endogenous stopping is strictly more permissive in some senses and simpler to summarize.
  • Practical consequence: If one source has strictly larger I0 and I1, then for any decision problem where states call for different optimal actions, F will eventually (for small enough c) deliver strictly higher value than G.

Data & Methods

  • This is a theoretical, mathematical paper; no empirical data are used.
  • Model setup:
    • Binary state θ ∈ {0,1}; finite, full-support, identified experiments F = (F0, F1) and G = (G0, G1).
    • Sequential sampling with stopping times that can depend on realized signals; policies can use state-independent randomization.
    • Exact simulation: given source F, construct stopping policies whose output distributions exactly match a target experiment H under both states.
  • Main technical tools and constructions:
    • Likelihood-ratio processes for stopped filtrations and martingale representations; use of KL divergence identities for stopped experiments (martingale optional stopping and KL decomposition).
    • A continuous-revelation embedding (using Brownian-bridge conditioning) to create continuous likelihood-ratio paths from discrete signal jumps so sequential stopping can hit required likelihood levels.
    • Two-point embedding and concatenation arguments to build simulators that achieve both statewise KL bounds simultaneously (up to the additive source-only constant).
    • Use of Jensen’s inequality to obtain the information lower bounds.
    • Small-cost sequential decision asymptotics (Chernoff/Wald-style analysis) to link KLs to value in costly-observation problems.
  • Assumptions and scope: finite signal spaces, full support, mutual absolute continuity for target laws, iid observations, binary state. The additive constant depends only on the source’s extreme likelihood-ratio values.

Implications for AI Economics

  • Metric for sequential usefulness of predictors/models: When predictions (signals) are acquired sequentially and observation cost is small, the two directed KL divergences (KL under each true state) are the right scalar summary to compare binary predictors/models. If a predictive model has strictly higher KL in both directions, it will dominate all others in sequential decision value (for small per-observation cost).
  • Model procurement and data collection choices: For decisions that allow adaptive sampling (stop when sufficient evidence accrued), decision-makers should prioritize sources/models with larger I0 and I1 rather than sources that dominate under a fixed-sample comparison metric (which may require considering Rényi divergences).
  • Active learning and labelling: In settings where labels/information can be acquired at low marginal cost and the learner can stop adaptively, the KL coordinates determine how many labels (in expectation) are needed under each state to reach a given informational target. This helps in budgeting and designing stopping criteria.
  • A/B testing and sequential evaluation: When comparing two predictive systems via sequential testing (with the ability to stop based on observed data), this framework clarifies when one system can reproduce the behavior of another with bounded additional sampling cost and when it will yield strictly higher decision value.
  • Policy and platform design: Platforms that mediate data or model selection (e.g., marketplaces for models or datasets) can use directed-KL-based metrics to rank sources for adaptive decision tasks. Contracts that price per observation should account that asymptotic value differences scale like c log(1/c)/Iθ(F), so small observation costs amplify the importance of KL differences.
  • Limitations and open directions relevant to AI economics:
    • Binary-state restriction: real-world decisions often involve many states or continuous parameters; extending to multi-state or continuous-state settings is necessary for broader applicability.
    • Finite-support and iid assumptions: predictors and data sources in AI may produce continuous, structured, or dependent signals; generalizing the conversion/simulation constructions to such settings is nontrivial.
    • Model misspecification and robustness: KL is sensitive to model misspecification; practical comparisons may need robustified divergences or worst-case criteria.
    • Strategic agents and costs: when observations or models are supplied by strategic parties (platforms/providers), incentives and pricing interact with information ordering — integrating this with the paper’s results is an important empirical-economic extension.

Takeaway for AI economics: when adaptive, low-cost sampling is feasible, compare candidate information sources/models by their directed KLs (statewise expected log-likelihood gains). This determines both feasibility of simulating other sources and asymptotic decision value ordering; more nuanced divergences (Rényi family) matter only when sample sizes are fixed a priori.

Assessment

Paper Typetheoretical Evidence Strengthhigh — The paper provides formal theorems with constructive proofs (conversion theorem, equivalence of stopping and decision dominance, and value consequences) and derives sharp statewise KL bounds and simulator constructions; results are mathematically rigorous within the stated assumptions. Methods Rigorhigh — The author develops precise definitions, proves necessary lower bounds (via likelihood-ratio/stopped σ-field arguments), gives a constructive simulator achieving the bounds up to an additive source-dependent constant, and uses standard, well-understood tools (KL divergences, martingale/stopping-time arguments, Brownian-bridge embedding) in a careful way. SampleNo empirical sample — the paper is purely theoretical: it analyzes finite, full-support, identified binary experiments with i.i.d. observations and derives properties of sequential simulators and decision problems with per-observation costs. Themesadoption productivity human_ai_collab GeneralizabilityRestricts attention to binary state problems (θ ∈ {0,1}); multistate extensions are nontrivial., Assumes finite signal spaces with full support and mutual absolute continuity; continuous or high-dimensional signals may require additional technical work., Assumes i.i.d. observations from the source experiment; dependent or nonstationary data are not covered., Key value-comparison conclusions use small-cost asymptotics (c ↓ 0); finite-cost behavior may differ., Constructive constants (additive terms) depend on the source experiment and may limit practical applicability to real-world high-dimensional AI outputs.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
For finite, full-support, identified binary experiments, stopping dominance and decision dominance are equivalent, and both hold if and only if the source experiment has weakly larger directed KL divergences in both states: I0(F) >= I0(G) and I1(F) >= I1(G). Decision Quality positive Value and informativeness of sequential experiments in finite decision problems
Reading fidelity high
Study strength high
I0(F) >= I0(G) and I1(F) >= I1(G)
0.2
Any exact simulation of a binary target experiment H using repeated observations from source experiment F must satisfy the statewise information lower bounds I0(H) <= I0(F) E0[N] and I1(H) <= I1(F) E1[N]. Task Completion Time negative Expected number of observations required for exact sequential simulation
Reading fidelity high
Study strength high
I0(H) <= I0(F)E0[N]; I1(H) <= I1(F)E1[N]
0.2
For every finite binary target experiment H, there exists a single exact simulator using source experiment F that simultaneously reaches both statewise KL lower bounds up to an additive constant depending only on F. Task Completion Time positive Efficiency of exact sequential experiment simulation
Reading fidelity high
Study strength high
I0(F)E0[N] <= I0(H) + log(r̄F/r_F); I1(F)E1[N] <= I1(H) + log(r̄F/r_F)
0.2
If F strictly dominates G under the sequential comparison order, then for every fixed finite decision problem, F yields at least as high a value as G for sufficiently small observation costs. Decision Quality positive Optimal expected utility in a sequential decision problem with costly observations
Reading fidelity high
Study strength high
Uc(F; µ, u) >= Uc(G; µ, u) for every sufficiently small c
0.2
When no single action is optimal in both states, strict dominance produces strictly higher decision value for F than for G once the observation cost is sufficiently small. Decision Quality positive Optimal expected utility from learning and acting in a finite decision problem
Reading fidelity high
Study strength high
Uc(F; µ, u) > Uc(G; µ, u) for every c in (0, c̄)
0.2
If the same action is optimal in both states, then the values of strictly dominant and dominated experiments are equal for every positive observation cost. Decision Quality null_result Optimal expected utility in a sequential decision problem with costly observations
Reading fidelity high
Study strength high
Uc(F; µ, u) = Uc(G; µ, u) for every c > 0
0.2
With fixed sample sizes, asymptotic experiment conversion is governed by the full spectrum of Rényi divergences, whereas allowing the sample size to depend on realized evidence reduces the relevant constraints to the two directed KL divergence ratios. Task Completion Time mixed Asymptotic number of source observations required to reproduce another experiment
Reading fidelity high
Study strength medium
maxθ∈{0,1} Iθ(G)/Iθ(F)
0.12

Notes