The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI models optimized for raw accuracy can be the wrong product: firms should choose how often models predict (coverage) to fit users’ verification costs and error losses, and small changes in those downstream economics can cause abrupt shifts in the optimal training target.

Training AI For When Humans Will Use It
Kevin A. Bryan, Joshua S. Gans · August 12, 2026
arxiv theoretical n/a evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Kevin A. Bryan unresolved corpus identity
  2. Joshua S. Gans unresolved corpus identity

Semantic Scholar

Latest observation:

  1. K. Bryan provider ID
  2. Joshua Gans provider ID
The economic value of an AI depends on how humans use its predictions, so optimal training trades off coverage and conditional correctness to match users' verification costs and loss asymmetries, meaning maximizing unconditional accuracy is generally suboptimal and optimal targets can change discontinuously when continuation plans switch.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI predicts; humans use its predictions to make decisions. These predictions are combined with human verification and analysis, queries to other statistical models, and so on. The economic value of an AI, therefore, depends on how it interacts with the surrounding decision environment. We describe the value of AI as part of this ``composite experiment'' where AI makes a coarse prediction of the state of the world, show what this means for optimal model training via a geometric argument, explain why optimal training can be discontinuous in economic variables, and study how heterogeneous users or monopoly model trainers affect these results. In particular, maximizing the unconditional accuracy of AI predictions is generally suboptimal.

Summary

Main Finding

The economic value of an AI prediction depends on how humans use that prediction inside a larger decision process (the “composite experiment”). Maximizing unconditional accuracy (overall percent correct) is generally not the right training objective. Instead, optimal model training trades off coverage (probability the model issues a prediction) against conditional correctness (probability the prediction is correct given the model predicts), and the downstream human continuation plan picks the tangency point on a model’s accuracy–coverage frontier. Because users may switch between qualitatively different continuation plans (e.g., verify vs trust), small changes in economic parameters can produce discontinuous jumps in the optimal training target.

Key Points

  • Basic primitives and notation
    • Model m = (q, p, c): coverage q = Pr(model predicts), conditional correctness p = Pr(prediction correct | model predicts), query cost c. Unconditional accuracy a = q p.
    • Agent payoffs: correct label gives H > 0, incorrect gives −L < 0, outside option normalized to 0.
  • Composite experiment and continuation plans
    • A continuation plan is the agent’s full contingent decision rule after observing the model (and possibly additional actions such as verification or consulting a second model).
    • Under a coverage-invariance assumption (the downstream continuation signals and costs conditional on coarse AI outcomes do not depend on (q,a)), the payoff from any continuation plan is affine in (q,a).
    • The agent’s value is the upper envelope (convex, piecewise-linear) of these affine plan payoffs and the baseline.
  • Verification example (pedagogical core)
    • With a perfect verification cost cV, the undominated downstream plans are: don't query; query and trust the model; query and verify (act only if verified).
    • Plan payoffs reduce to affine forms like Π(q,a) = α + β a − γ q − c, with slopes determined by H, L, cV.
  • Training frontier and tangency principle
    • Training yields a frontier of feasible (q,p) (or equivalently of attainable unconditional accuracy a as a function of q). Typically assumed concave: attaining higher coverage requires accepting lower conditional correctness.
    • For a given continuation plan (an affine payoff), optimal training is found at the tangency between that plan’s iso-payoff line and the accuracy–coverage frontier.
    • When the agent’s optimal continuation plan changes (e.g., verification becomes too costly and the user switches from verify→trust), the tangency can jump to a different point on the frontier → discontinuous changes in the optimal (q,p) target.
  • Consequences and comparative statics
    • Environments with cheap verification (or low downside L) favor high coverage: catch more cases and verify cheaply.
    • Environments with expensive verification and high L favor selective models with low coverage but high conditional correctness.
    • Heterogeneous users served by a single trained model create welfare losses; monopolists optimally target marginal users rather than population-average welfare and may strategically offer coarser outputs (e.g., subscription tiers).
  • Connections to ML literature
    • Relates to selective classification / reject option and learning-to-defer literatures; difference: this paper takes the continuation technology and agent payoffs as primitives and asks where on the training frontier (q,p) the model should operate.
    • Also relates to Blackwell comparisons of experiments, but Blackwell orderings need not hold here because of coarse outputs and the coverage–accuracy tradeoff.
  • Robustness / scope
    • Main results hold for a broad class of finite, coverage-invariant continuation technologies (multiple verification strategies, partially correlated alternative models, etc.). Infinite policy classes need extra regularity.
    • Baseline assumes model outputs are coarse (predict label or abstain); allowing calibrated case-specific posteriors changes the trainer’s problem (possible to reveal quality instead of choosing coverage).

Data & Methods

  • Methodology: analytic, theoretical model and comparative-statics. No empirical estimation or dataset used.
  • Core formal steps:
    • Define primitives (q, p, a, H, L, costs) and the agent’s sequential decision problem (composite experiment).
    • Show continuation payoffs Πr(q,a) are affine under coverage-invariance (Lemma/Proposition).
    • Define a model training frontier mapping feasible conditional correctness to coverage (concave frontier commonly assumed).
    • Use geometric/tangency arguments: for each affine continuation payoff, the optimal (q,a) solves a tangency between an iso-payoff line and the frontier; agent’s value is the upper envelope over plans.
    • Analyze comparative statics and show possible discontinuities when the agent switches continuation plans.
    • Extend framework to consider heterogeneous users, monopolist model trainer incentives, subscription/coarse-output strategies, team verification, adversarial settings.
  • Key assumptions to note:
    • Coverage-invariance: downstream signals and costs, conditional on coarse events (no-prediction, correct, wrong), do not change with the model’s global (q,a).
    • Coarse outputs: model gives a label or abstains rather than calibrated posteriors (baseline).
    • Risk-neutral agents; finite set of continuation plans (for formal piecewise-linear envelope argument).

Implications for AI Economics

  • For model designers and firms
    • Training objectives must incorporate downstream human behavior and verification costs. Maximizing unconditional accuracy is not generally welfare-maximizing.
    • Decide whether to prioritize coverage (more predictions, lower per-prediction accuracy) or selectivity (abstain on hard cases): criterion depends on H, L, verification costs, and the human continuation technology.
    • Firms (especially monopolists) may strategically choose model coarseness or coverage to target marginal consumers and may offer tiered/coarse outputs for price discrimination.
  • For practitioners and procurement/evaluation
    • Evaluate models by their value inside realistic composite experiments (how will users verify, defer, or trust?), not only by global accuracy metrics.
    • When deploying predictive systems in domains with cheap verification (e.g., cheap tests, code testing), favor broader-coverage models; where verification is costly and errors are costly (e.g., high-stakes medical/safety decisions), favor selective models or provide calibrated uncertainty.
    • Be alert to discontinuities: small changes in verification costs or loss parameters (or policy/regulation) can justify materially different training targets and deployment rules.
  • For policy and regulation
    • Standards and procurement specifications should reflect the downstream decision environment (verification capacity, error costs) rather than raw accuracy comparisons.
    • Encouraging or enabling verification capacity (e.g., cheaper confirmatory tests, better human-AI workflows) can change the societal optimal design of AI systems toward higher coverage and broader usefulness.
  • For research
    • Empirical work: measure verification costs, downstream loss magnitudes, and how users switch between continuation plans in real settings—these parameters determine the optimal modeling tradeoffs.
    • Theoretical extensions: relax coverage-invariance, allow calibrated/posterior outputs, endogenize verification technology investments, and consider infinite (or continuous) policy classes for richer applications.
    • ML methodology: integrate downstream payoff structure into training objectives or selection among selective classifiers / abstention rules; study methods to reveal case-level uncertainty (calibration) so that coverage need not be hard-coded by the trainer.

Short takeaway: pick the training target (coverage vs accuracy) with the human decision environment in mind. The best AI for welfare is the one that, together with user verification and follow-up options, produces the best composite experiment — not necessarily the one with the highest raw accuracy.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is purely theoretical and provides formal propositions and geometric proofs rather than empirical or experimental evidence; thus empirical strength is not applicable. Methods Rigorhigh — Theoretical development is clear and internally consistent: payoffs are decomposed, continuation plans shown to be affine under coverage-invariance, and optimal training characterized via a tangency principle on a well-defined accuracy-coverage frontier; proofs appear to handle extensions (heterogeneity, monopoly, adversaries) and relate to existing literatures. Rigor is reduced somewhat by simplifying assumptions (coarse outputs/abstain only, coverage-invariance, finite set of continuation plans, risk-neutrality, perfect verification), which limit applicability without further empirical calibration. SampleNo empirical sample; the paper uses a formal model where the AI is characterized by coverage q and conditional correctness p (unconditional accuracy a = qp), agents choose continuation plans (trust, verify, abstain) given parameters H (gain from correct action), L (loss from wrong action), query cost c and verification cost cV, and a pre-specified concave training frontier mapping feasible (q,p) pairs. Extensions consider heterogeneous users, monopoly trainers, and other continuation technologies. Themeshuman_ai_collab productivity org_design adoption IdentificationAnalytical theoretical model: derives causal implications through a formal decision-theory framework (composite experiments), algebraic decomposition of continuation payoffs, and geometric tangency arguments on an accuracy-coverage training frontier under stated assumptions (coverage-invariance, finite continuation plans, risk-neutral agents). No empirical identification or data-based causal inference. GeneralizabilityAssumes coarse AI output (predict-or-abstain) rather than calibrated case-specific posteriors; limits direct mapping to systems that provide rich uncertainty estimates., Coverage-invariance assumption: downstream verification costs and signal distributions are assumed independent of (q,a); real-world marginal cases may violate this., Finite set of continuation plans and one-shot decision setting; dynamic, repeated, or continuous policy classes are not fully analyzed., Risk-neutral agents assumed; risk aversion could alter tangency conditions and optimal coverage., Perfect verification (in main examples) is assumed; imperfect or noisy verification would change payoffs., Model is not empirically estimated or validated against field data; quantitative magnitudes and frictional features are not calibrated., Strategic interactions and market competition are treated in reduced form; richer incentive or multi-agent dynamics may change trainer objectives.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Maximizing the unconditional accuracy of an AI model is generally not the economically optimal training objective when humans use the model’s predictions as inputs into further decisions. Decision Quality negative Economic value and downstream decision performance of the AI model
Reading fidelity high
Study strength high
not reported
0.2
The economically relevant object is a composite experiment consisting of the AI model, the human user, and the user’s available follow-up actions, rather than the AI prediction alone. Decision Quality positive Economic value generated by integrating AI predictions with human decision processes
Reading fidelity high
Study strength high
not reported
0.2
Users trust an AI prediction without verification only when verification is sufficiently expensive relative to the expected loss from an erroneous action. Task Allocation positive User adoption and verification behavior following an AI prediction
Reading fidelity high
Study strength high
not reported
0.2
When both mistakes are very costly and verification is expensive, users may choose not to use the AI at all. Adoption Rate negative Whether the AI is used in the decision process
Reading fidelity high
Study strength high
not reported
0.2
With a concave accuracy-coverage training frontier, the optimal training target is determined by a tangency between the frontier and an iso-payoff line whose slope depends on the user’s downstream continuation plan. Decision Quality positive Optimal AI training target along the coverage-accuracy frontier
Reading fidelity high
Study strength high
not reported
0.2
Optimal AI training can change discontinuously as human decision parameters change, even when the value of the best available decision plan remains continuous. Task Allocation mixed Optimal model coverage and conditional correctness as verification costs and error losses vary
Reading fidelity high
Study strength high
not reported
0.2
Under coverage-invariant continuation technologies, every continuation plan has an affine payoff in coverage and unconditional accuracy. Decision Quality positive Expected payoff from a human-AI continuation plan
Reading fidelity high
Study strength high
not reported
0.2
When the set of continuation plans is finite and continuation technologies are coverage-invariant, the user’s value function is the upper envelope of affine plan payoffs and is therefore convex and piecewise linear in coverage and unconditional accuracy. Decision Quality positive User value as a function of AI coverage and unconditional accuracy
Reading fidelity high
Study strength high
not reported
0.2
The paper’s training and discontinuity results extend beyond the trust-or-verify example to finite composite decision problems involving additional information, partially correlated alternative AI models, and other sequential continuation technologies, provided coverage-invariance holds. Organizational Efficiency positive Applicability of the model-training principle across human-AI decision environments
Reading fidelity high
Study strength medium
not reported
0.12

Notes