The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new embedding theorem collapses dynamic exploration-and-stopping problems into a static convex program priced by a shadow value of information; this yields that convex time preference induces Poisson-style exploration, concave preference produces delayed pure-exploration windows, and the linear case yields Brownian exploration.

Exploration and Stopping
Yuliy Sannikov, Weijie Zhong · August 10, 2026
arxiv theoretical n/a evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yuliy Sannikov unresolved corpus identity
  2. Weijie Zhong unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Yuliy Sannikov provider ID
  2. Weijie Zhong provider ID
The paper proves that any admissible exploration-and-stopping strategy is equivalent to choosing a joint distribution over stopped states and stopping times subject to per-date information-budget inequalities, reducing the dynamic problem to a static convex program whose dual prices information over time and whose solutions imply Poisson, Brownian or pure-exploration regimes depending on the curvature of time preference.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We study a decision-maker who explores --- dynamically choosing what to learn --- before stopping to act. We first reduce this dynamic control problem to a static one: any exploration-and-stopping strategy is equivalent to a choice of the joint distribution of the stopped state and the stopping time, subject to one information-budget constraint at each date, and we characterize exactly which distributions are attainable. The reduced problem is a convex program with a linear objective; its dual prices information over time, and the optimal policy concavifies the stopping payoff net of these shadow prices. The curvature of the decision-maker's time preference then governs the shape of optimal exploration: convex time preference induces Poisson exploration, concave time preference confines stopping to a window whose length is controlled by the dispersion of the marginal cost of delay --- forcing an initial phase of pure exploration when the window is short --- and the linear case lies at the boundary between them. We apply the framework to real options, to the speed--accuracy tradeoff in information acquisition, and to a continuous-time exploration contest.

Summary

Main Finding

The paper gives a complete reduction of dynamic exploration-and-stopping problems (where the decision-maker chooses what information to acquire over time and when to stop) to a static convex program over joint distributions of stopped state and stopping time. It provides an exact, date-by-date characterization (an "embedding theorem") of which state–time distributions are attainable under a single information-rate constraint, derives the dual shadow-price path that prices information over time, and shows that the curvature of time preference determines the qualitative form of optimal exploration (Poisson-like, Brownian-like, or an initial pure-exploration window). Several applications illustrate implications for real options, the speed–accuracy tradeoff, and continuous-time exploration contests.

Key Points

  • Setup and admissibility:

    • The controlled process is a martingale µt taking values in a convex compact S; exploration choices determine the martingale law.
    • Information flow is summarized by a convex information measure H and a learning-rate parameter χ. Admissible exploration processes satisfy an expected H-variation bound: E[H(µt′) − H(µt) | Ft] ≤ χ (t′ − t) for t′ > t.
    • The DM also chooses a stopping time τ. The object of interest is the joint law f of (µτ, τ).
  • Embedding theorem (Theorem 1):

    • A distribution f over S × T is attainable if and only if 1) Ef[µ] = prior µ0 (martingale/mean constraint), and 2) for every date t, an information-budget inequality holds: ∫{τ≤t} H(µ) f(dµ,dτ) + H(∫{τ>t} µ f(dµ,dτ)) − H(µ0) ≤ χ · E[min{t, τ}].
    • Intuition: at each date t, the information already exploited (mass stopped by t) plus the dispersion still held in reserve (the continuation atom) cannot exceed the budget accumulated up to t.
  • Reduction to a static convex program:

    • The set of feasible state–time laws F is convex and characterized by the mean constraint and one concave constraint per date → the exploration-and-stopping control problem reduces to choosing f ∈ F to maximize a linear objective (expected payoff).
    • Because the reduced problem is a convex program with linear objective, strong duality applies.
  • Duality and shadow prices:

    • The date-by-date constraints have multipliers that aggregate into a nonincreasing shadow-price path Λ(t), the marginal value of relaxing the information budget at time t.
    • Complementary slackness: Λ declines only when the information constraint binds; where slack, Λ is flat.
    • Optimal f concavifies the (stopping) payoff augmented by the shadow-value; equivalently, mass is placed on the upper tangent to payoff + shadow value.
    • Solving the optimal policy reduces to solving a one-dimensional integral/ODE for Λ(t) in regular cases. The paper gives two algorithms to compute Λ(t).
  • Trichotomy by curvature of time preference:

    • Convex adjusted time preference (time-risk-loving) → Poisson exploration: the state moves deterministically but at a random (Poisson) arrival jumps into stopping region.
    • Linear time preference → Brownian-type (knife-edge) exploration.
    • Concave time preference (time-risk-averse) → stopping confined to a finite window; when the window is short, optimal policy begins with pure exploration (belief dispersion with no stopping) followed later by exploitation.
    • This unifies and generalizes prior special-case results (e.g., exponential discounting, binary threshold rules).
  • Applications:

    • Real options with active exploration: recovers and extends HJB-style results; working example throughout the paper.
    • Speed–accuracy tradeoff: within an interior regime, optimal policies alternate between pure exploration (belief disperses; no stopping) and full exploitation (jumps and immediate stopping); relates curvature of adjusted discounting to whether decision quality increases or decreases with response time — providing a normative link with empirical latency–accuracy patterns.
    • Competitive exploration contest: when n players privately choose exploration processes and the first to stop wins a quality-dependent prize, all nontrivial symmetric pure-strategy equilibria use Poisson exploration; competition endogenously induces an effectively convex (time-risk-loving) discounting.

Data & Methods

  • Nature of contribution: fully theoretical / mathematical. No empirical dataset is used.
  • Mathematical objects and assumptions:
    • Continuous- or discrete-time horizon T (contains 0 and at least one other point).
    • State µt: cadlag martingale in compact convex S ⊂ R^n, starting at µ0.
    • Information measure H: strictly convex, continuous on S, extended homogeneously off S; learning-rate χ > 0.
    • Admissible pairs (µt, τ) satisfy the conditional H-variation bound for all t′ > t.
  • Core methods:
    • Martingale embedding / realization: characterize attainable joint laws of stopped state and time when both the martingale law and stopping time are chosen subject to the H-variation clock.
    • Convex analysis and linear programming in infinite dimensions: reduce dynamic control to a static convex program over measures, exploit linearity of objective to derive strong duality.
    • Duality and complementary slackness: interpret dual multipliers as time-dependent information shadow prices Λ(t); use concavification arguments (mass placed on upper tangents) to derive structure of optimal f.
    • Construction/proof techniques: optional sampling, convex-order (peacock) arguments (Strassen–Kellerer tradition), explicit realization of martingale paths to attain any feasible f, ODE/integral characterization of Λ(t) in regular cases.
  • Examples and specializations:
    • H quadratic → recovers diffusive/quadratic-variation bounded learning (Brownian).
    • H = negative entropy → mutual-information-rate constraint (Poisson/shot-noise learning allowed).
    • The paper includes two algorithms for computing Λ(t) numerically (details in Section 3).
  • Formal deliverable: Theorem 1 (characterization of embeddable f), solution of the primal/dual convex program, classification of optimal exploration by time-preference curvature, and worked applications.

Implications for AI Economics

  • Conceptual implications
    • Explaining exploration vs deployment tradeoffs: the paper gives a compact, computationally tractable way to derive optimal exploration policies when an AI developer chooses what signals/tests to perform before deploying a model or product, subject to an information-rate constraint (resource, compute, lab time).
    • One-step design: modelers can optimize over distributions of stopping times and final beliefs rather than infinite-dimensional signal/control paths, simplifying comparative statics and mechanism design.
    • Value-of-information path: the dual Λ(t) provides a principled, time-resolved shadow price of information — useful for pricing data acquisition, compute, or experimentation capacity, and for deciding dynamic subsidies or taxes.
  • Predictions and testable empirical regularities
    • Curvature of time preference (including effective curvature induced by competition or prizes) predicts the qualitative learning process:
      • Competition or incentives that convexify effective time preference → Poisson-like, front-loaded random stopping (risk-seeking over time).
      • Institutional features that imply concavity (risk-averse over timing) → initial pure-exploration phases with delayed stopping and stopping concentrated in a later window.
    • Speed–accuracy patterns: monotone relationships between response time and decision quality are driven by curvature of adjusted discounting; exponential (linear) discounting implies time-invariant quality, convex implies quality decreases with delay, concave implies quality increases initially — this links lab/field measurements of latency–accuracy curves to implied time-preference shapes and information costs.
  • Applications to AI R&D and markets
    • Firm experimentation and launch timing: for AI firms deciding test suites and deployment timing in the presence of competitor entry, the framework predicts when firms will do short, intense (Poisson-like) testing vs long accumulation phases.
    • Contest and prize design: tournament or prize structures that alter effective time-preference curvature can be used to steer firms toward faster (Poisson) or more thorough (pure-exploration) research processes.
    • Mechanism design for information markets: Λ(t) gives a target for designing dynamic pricing of compute or labeled data over time to induce desired exploration profiles.
  • Practical computational implication
    • The reduction to a convex program (linear objective) and a one-dimensional ODE/integral for shadow prices suggests implementable numerical approaches for computing optimal exploration policies in applied AI-economic models (e.g., platforms allocating testing budgets, regulators pricing audits, firms planning phased model evaluations).
  • Limitations and caveats
    • The theory is abstract and relies on the H-variation bound and martingale representation; real-world frictions (nonmartingale information, strategic disclosure rules, multi-stage contracts, unmodeled costs) may complicate direct application.
    • Behavioral or institutional factors that set time-preference curvature are exogenous in the baseline model (though the contest application shows endogenous convexification).
    • Extensions (some discussed in the paper) allow endogenous capacity or richer constraints, but empirical calibration will require structural choices for H and χ.
  • Suggestions for AI-economics researchers
    • Use the embedding reduction to replace complex dynamic signal-design problems with finite-dimensional convex programs when studying deployment timing, model validation, or R&D scheduling under information limits.
    • Estimate the implied curvature of effective time preference from observed latency–accuracy data to infer underlying exploration incentives or information costs.
    • In policy or prize design, target the effective Λ(t) profile (via subsidies, disclosure rules, or prize timing) to steer private exploration toward socially preferred timing/quality tradeoffs.

Assessment

Paper Typetheoretical Evidence Strengthn/a — Purely theoretical/mathematical contribution: the paper develops and proves an embedding theorem and derives comparative-static implications; there is no empirical identification or causal estimation to evaluate. Methods Rigorhigh — The paper states and proves a general embedding theorem, reduces a dynamic control problem to a finite-dimensional convex program, uses duality to characterize shadow prices, and derives sharp comparative-static results; arguments are grounded in established martingale, convex-order, and optimal-transport literatures and connect to multiple prior results as special cases. SampleNo empirical sample — a continuous-time (or discrete-time) theoretical model: beliefs µ_t form a martingale on a convex compact state space S; information acquisition constrained by a convex information measure H and rate parameter χ; admissible strategies are cadlag martingales satisfying an H-variation bound; object of study is the set of joint distributions over stopped state µ_τ and stopping time τ and the optimal choice of such distributions. Themesinnovation human_ai_collab GeneralizabilityAbstract, stylized continuous-time martingale model may not map directly to discrete, institutional, or operational information processes used by firms or AI systems, Assumes the decision-maker can freely design the information process subject only to an H-variation bound; real-world signal design or data-collection frictions may limit feasibility, Key primitives (choice of H, rate χ, state space S, prior µ0, time-preference) must be calibrated for applied settings; results depend on those choices, Strategic extensions (multiple heterogeneous agents, asymmetric information, complex contracting, market frictions) may alter equilibrium or optimal policies beyond the classes analyzed, Some comparative predictions rely on regularity (e.g., differentiability) and convexity/concavity assumptions that may fail in specific applications

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Any exploration-and-stopping strategy can be reduced to a choice of the joint distribution of the stopped state and stopping time, subject to one information-budget constraint at each date. Organizational Efficiency positive Complexity of solving the exploration-and-stopping problem
Reading fidelity high
Study strength high
not reported
0.2
The attainable state-time distributions are characterized exactly by a mean-preservation condition and a date-by-date information-budget inequality. Task Allocation positive Attainability of joint stopped-state and stopping-time distributions
Reading fidelity high
Study strength high
not reported
0.2
The reduced exploration-and-stopping problem is a convex program with a linear objective, and its dual assigns shadow prices to information over time. Organizational Efficiency positive Computational tractability and pricing of information constraints
Reading fidelity high
Study strength high
not reported
0.2
When adjusted time preference is convex, optimal exploration resembles Poisson exploration: the state follows a deterministic path and then jumps into the stopping region at a random time. Task Allocation positive Optimal exploration pattern under convex time preference
Reading fidelity high
Study strength high
not reported
0.2
When adjusted time preference is concave, optimal stopping is confined to a window whose length is governed by the dispersion of the marginal cost of delay; if the window is short, the policy begins with a phase of pure exploration. Task Allocation positive Optimal exploration and stopping pattern under concave time preference
Reading fidelity high
Study strength high
not reported
0.2
The linear time-preference case is the boundary between the convex and concave regimes, with Brownian exploration being optimal. Task Allocation mixed Optimal exploration process under linear time preference
Reading fidelity high
Study strength medium
not reported
0.12
In the paper's speed-accuracy application, the curvature of adjusted time preference determines whether decision quality rises or falls with response time. Decision Quality mixed Decision quality as a function of response time
Reading fidelity high
Study strength medium
not reported
0.12
In the speed-accuracy application, optimal behavior alternates between pure exploration and full exploitation: under pure exploration beliefs disperse without stopping, while under full exploitation the decision-maker stops when the belief jumps. Task Allocation mixed Allocation between information acquisition and stopping
Reading fidelity high
Study strength medium
not reported
0.12
In the continuous-time exploration contest, every nondegenerate state-exchangeable pure-strategy equilibrium is player-symmetric and uses Poisson exploration. Task Allocation positive Equilibrium exploration strategy in a competitive R&D contest
Reading fidelity high
Study strength medium
not reported
0.12
Competition endogenously convexifies the effective discount factor, thereby generating Poisson exploration in nondegenerate equilibria. Task Allocation positive Effect of competition on exploration dynamics
Reading fidelity high
Study strength medium
not reported
0.12

Notes