The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI 'speedup' claims are often non‑comparable and misleading because they ignore quality tradeoffs, rework, and integration costs; the authors propose measuring productivity as Time‑To‑Acceptance under documented acceptance tests. They introduce a quantitative framework and a lightweight Human‑AI Productivity Card to normalize task complexity, capture uncertainty and rework, and standardize reporting across studies.

Human-AI productivity claims should be reported as time-to-acceptance under explicit acceptance tests
Chaoyue He, Xin Zhou, Di Wang, Hong Xu, Wei Liu, Chunyan Miao · February 06, 2026
openalex commentary n/a evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Chaoyue He provider ID
  2. Xin Zhou provider ID
  3. Di Wang provider ID
  4. Hong Xu provider ID
  5. Wei Liu provider ID
  6. Chunyan Miao provider ID
The paper argues that claims about human‑AI 'speedups' should be standardized and reported as Time‑To‑Acceptance under documented acceptance tests with complexity normalization, explicit treatment of uncertainty and rework, and recommends a Human‑AI Productivity Card to make productivity evidence cumulative and comparable.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This position paper argues that human-AI productivity claims should be treated as scientific claims requiring standardized, complexity-normalized, uncertainty-aware reporting rather than ad hoc "speedup" numbers. In the pre-AGI transition, general-purpose AI is becoming cognitive infrastructure across tasks yet remains unevenly reliable, producing jagged performance frontiers and non-comparable results across studies. Reported gains can silently trade quality for speed and often omit verification, rework, integration, and governance costs that dominate real deployments. We argue that productivity should be reported as Time-To-Acceptance (TTA) under a documented acceptance test, with task complexity held invariant (or explicitly normalized) and heterogeneity and uncertainty treated as first-class quantities. We propose a quantitative framework that reallocates (rather than erases) complexity across human/AI workflow stages, makes rework explicit via acceptance probabilities, and captures non-linear scaling through effective parallelism and collaboration efficiency. Finally, we introduce the Human-AI Productivity Card (HAP Card), a lightweight reporting artifact to support cumulative and comparable evidence on human-AI collaboration in the pre-AGI transition.

Summary

Main Finding

Productivity claims for human-AI collaboration should be treated as scientific claims and reported using standardized, complexity-normalized, uncertainty-aware measures rather than ad hoc "speedup" numbers. The paper proposes Time‑To‑Acceptance (TTA) under a documented acceptance test as the primary metric, together with a modeling framework that (a) reallocates task complexity across human/AI workflow stages, (b) makes rework explicit via acceptance probabilities, and (c) captures non‑linear scaling through effective parallelism and collaboration efficiency. To operationalize this, the authors introduce the Human‑AI Productivity Card (HAP Card): a lightweight, standardized reporting artifact to enable cumulative, comparable evidence during the pre‑AGI transition.

Key Points

  • Current practice: reported speedups are often incomparable and misleading because they
    • do not hold task complexity constant,
    • trade quality for speed without explicit verification,
    • omit rework, integration, verification, and governance costs that dominate deployment, and
    • gloss over heterogeneity and uncertainty (task instances, workers, model stochasticity).
  • Pre‑AGI context: general‑purpose AI is becoming cognitive infrastructure but is unevenly reliable, producing jagged, non‑monotonic performance frontiers across tasks and settings.
  • Proposal: report productivity as Time‑To‑Acceptance (TTA):
    • TTA is the expected wall‑clock time until a task output meets a pre‑specified acceptance test.
    • Task complexity must be held invariant or explicitly normalized across comparisons.
    • Heterogeneity (across tasks, workers, instances) and uncertainty (sampling, model stochasticity, nonstationarity) must be reported as first‑class quantities (distributions, confidence intervals).
  • Rework explicitness: model rework with an acceptance probability p and explicit times for production, verification, and fixing; expected TTA increases with lower p and larger rework costs.
  • Nonlinear scaling: introduce metrics like effective parallelism (how much parallel work actually reduces TTA) and collaboration efficiency (how well human and AI labor combine), since adding workers or models has diminishing and task‑dependent returns.
  • Human‑AI Productivity Card (HAP Card): a standardized template listing task spec, acceptance test, time components, acceptance probabilities, uncertainty, model and prompt details, and costs to enable meta‑analysis and reproducibility.

Data & Methods (recommended framework)

  • Primary metric
    • Time‑To‑Acceptance (TTA): expected time from task start to an output that passes a documented acceptance test.
    • Report mean/median and full distribution (or percentiles) and confidence intervals.
  • Model of rework (simple representation)
    • Let p = probability a candidate is accepted on an iteration.
    • Let t_prod = time to produce a candidate, t_ver = verification time, t_fix = time to fix a rejected candidate.
    • Expected TTA (simple form): E[TTA] = ((1 − p)/p) * (t_prod + t_ver + t_fix) + (t_prod + t_ver). (Derivable from geometric failures; report component breakdowns.)
    • Report p and component times separately; sensitivity analyses when p depends on task instance or model version.
  • Complexity normalization
    • Hold task specification (input distribution, instructions, acceptance test) invariant across comparisons.
    • When tasks differ, use explicit complexity normalization (e.g., information/content measures, standardized task buckets, cognitive work components) and report normalized metrics.
  • Scaling and collaboration metrics
    • Effective parallelism = realized speedup from adding parallel agents relative to ideal linear parallelism.
    • Collaboration efficiency = ratio of achieved productivity to sum of isolated agent productivity (captures interaction overheads).
    • Report how these metrics vary with team size, model scale/version, and task structure.
  • Uncertainty & heterogeneity
    • Report within‑task and across‑task variance, worker heterogeneity, model stochasticity, model drift (versioning), and sensitivity to prompts/data.
    • Use randomized controlled designs where possible; pre‑register acceptance tests and analysis plans.
  • Costs & externalities
    • Report verification, rework, engineering/integration, monitoring, governance and compliance costs separately from raw generation time.
    • Convert to economic measures when relevant (e.g., labor hours per accepted unit, $ cost per accepted unit).
  • HAP Card (recommended fields)
    • Task name and exact specification (dataset, input distribution)
    • Acceptance test: exact rubric/automation used
    • Model(s) & version, prompt/chain-of-thought/template used
    • TTA: mean/median, percentiles, CI
    • Component times: t_prod, t_ver, t_fix (distributions)
    • Acceptance probability p (with CI)
    • Effective parallelism and collaboration efficiency (with measurement protocol)
    • Sample sizes, worker population, data splits
    • Costs: time→$ breakdown (integration, monitoring, governance)
    • Sensitivity analyses & caveats
    • Links to data, code, and pre‑registered protocol
  • Empirical design suggestions
    • Compare human‑only, AI‑only, and hybrid workflows on identical task specs.
    • Randomize across task instances and workers; report distributional effects (not just averages).
    • Report model versioning and repeat key tests over time to capture nonstationarity.

Implications for AI Economics

  • Measurement quality and growth accounting
    • Standardized TTA and HAP Cards will improve estimates of productivity gains attributable to AI and avoid over‑estimating speedups that ignore quality/rework costs.
    • Better microdata for growth accounting: more accurate contributions of AI to labor productivity and total factor productivity.
  • Labor market modeling
    • Explicit modeling of acceptance probabilities and rework clarifies which tasks are automatable and which require persistent human oversight, improving substitution vs. complementarity estimates.
    • Heterogeneity reporting allows richer models of occupational impacts and wage dynamics by capturing variation in task complexity and worker skills.
  • Investment and adoption forecasting
    • Decision‑makers and investors will get more realistic expected time‑to‑value numbers, including integration and governance overheads, reducing misallocation driven by optimistic, non‑comparable speedup claims.
  • Diffusion and scaling
    • Metrics for effective parallelism and collaboration efficiency inform models of how productivity scales with deployment size and team composition, affecting industry structure forecasts.
  • Policy and regulation
    • Standard reporting helps procurement, safety, and regulatory oversight by tying claims to an explicit acceptance test and documented uncertainty; enables audits and comparability across vendors.
  • Research & meta‑analysis
    • HAP Cards create a corpus of comparable measurements to enable meta‑analyses and evidence accumulation during the pre‑AGI transition, improving the scientific basis for economic modeling and policy.

Summary: Treat human‑AI productivity as a measured, uncertainty‑aware scientific claim. Use TTA under a documented acceptance test, explicitly model rework and scaling, normalize for task complexity, and adopt HAP Cards to make cross‑study comparisons and economic inferences robust.

Assessment

Paper Typecommentary Evidence Strengthn/a — This is a position/standards paper that does not present primary empirical identification or causal estimation; it proposes measurement conventions and a conceptual framework rather than empirical tests, so there is no empirical evidence to rate. Methods Rigorn/a — No empirical methods or data analysis are used; the contribution is conceptual (definition of Time-To-Acceptance, rework/acceptance-probability modeling, and the HAP Card). Rigor therefore pertains to clarity and internal consistency of the proposed framework rather than statistical methods. SampleNo primary data or sample; a conceptual position paper drawing on examples and prior literature to motivate a quantitative reporting framework and proposing a lightweight reporting artifact (the Human-AI Productivity Card) for future empirical work. Themesproductivity human_ai_collab org_design adoption governance GeneralizabilityNot empirically validated — framework not yet tested across tasks, industries, or datasets, Assumes feasibility of standardized acceptance tests, which may be difficult across domains with subjective quality criteria, May not capture domain-specific integration, regulatory, or organizational constraints (e.g., healthcare, legal, regulated industries), Focuses on pre-AGI general-purpose AI; may not apply to highly specialized automation or future AGI contexts, Implementation burden and measurement costs could limit adoption in small firms or low-resource settings, Cross-cultural and labor-institution differences (work practices, inspection norms) are not addressed

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Human-AI productivity claims should be treated as scientific claims requiring standardized, complexity-normalized, uncertainty-aware reporting rather than ad hoc "speedup" numbers. Research Productivity positive quality of reporting on human-AI productivity claims (proposed change in methodology)
Reading fidelity high
Study strength speculative
not reported
0.01
In the pre-AGI transition, general-purpose AI is becoming cognitive infrastructure across tasks. Automation Exposure positive extent to which general-purpose AI functions as underlying cognitive support across tasks
Reading fidelity high
Study strength low
not reported
0.03
General-purpose AI remains unevenly reliable, producing jagged performance frontiers and non-comparable results across studies. Output Quality negative reliability and comparability of AI performance across tasks/studies
Reading fidelity high
Study strength low
not reported
0.03
Reported gains can silently trade quality for speed and often omit verification, rework, integration, and governance costs that dominate real deployments. Output Quality negative tradeoff between speed (reported gains) and output quality plus omitted deployment costs
Reading fidelity high
Study strength low
not reported
0.03
Productivity should be reported as Time-To-Acceptance (TTA) under a documented acceptance test, with task complexity held invariant (or explicitly normalized) and heterogeneity and uncertainty treated as first-class quantities. Task Completion Time positive time until an output meets a documented acceptance test (Time-To-Acceptance)
Reading fidelity high
Study strength speculative
not reported
0.01
The paper proposes a quantitative framework that reallocates (rather than erases) complexity across human/AI workflow stages. Task Allocation neutral distribution of task complexity across human and AI workflow stages
Reading fidelity high
Study strength speculative
not reported
0.01
The framework makes rework explicit via acceptance probabilities. Task Completion Time neutral probability of initial output acceptance and consequent rework requirements
Reading fidelity high
Study strength speculative
not reported
0.01
The framework captures non-linear scaling through effective parallelism and collaboration efficiency. Team Performance neutral how productivity scales non-linearly with parallelism and collaboration efficiency
Reading fidelity high
Study strength speculative
not reported
0.01
We introduce the Human-AI Productivity Card (HAP Card), a lightweight reporting artifact to support cumulative and comparable evidence on human-AI collaboration in the pre-AGI transition. Research Productivity positive standardized reporting of human-AI productivity studies (HAP Card adoption and information captured)
Reading fidelity high
Study strength speculative
not reported
0.01
Adopting TTA, complexity normalization, explicit heterogeneity/uncertainty reporting, and the HAP Card will enable cumulative and comparable evidence on human-AI collaboration during the pre-AGI transition. Adoption Rate positive ability to accumulate comparable evidence across human-AI studies
Reading fidelity high
Study strength speculative
not reported
0.01

Notes