0 cumulative citations
View corpus contextAI 'speedup' claims are often non‑comparable and misleading because they ignore quality tradeoffs, rework, and integration costs; the authors propose measuring productivity as Time‑To‑Acceptance under documented acceptance tests. They introduce a quantitative framework and a lightweight Human‑AI Productivity Card to normalize task complexity, capture uncertainty and rework, and standardize reporting across studies.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This position paper argues that human-AI productivity claims should be treated as scientific claims requiring standardized, complexity-normalized, uncertainty-aware reporting rather than ad hoc "speedup" numbers. In the pre-AGI transition, general-purpose AI is becoming cognitive infrastructure across tasks yet remains unevenly reliable, producing jagged performance frontiers and non-comparable results across studies. Reported gains can silently trade quality for speed and often omit verification, rework, integration, and governance costs that dominate real deployments. We argue that productivity should be reported as Time-To-Acceptance (TTA) under a documented acceptance test, with task complexity held invariant (or explicitly normalized) and heterogeneity and uncertainty treated as first-class quantities. We propose a quantitative framework that reallocates (rather than erases) complexity across human/AI workflow stages, makes rework explicit via acceptance probabilities, and captures non-linear scaling through effective parallelism and collaboration efficiency. Finally, we introduce the Human-AI Productivity Card (HAP Card), a lightweight reporting artifact to support cumulative and comparable evidence on human-AI collaboration in the pre-AGI transition.
Summary
Main Finding
Productivity claims for human-AI collaboration should be treated as scientific claims and reported using standardized, complexity-normalized, uncertainty-aware measures rather than ad hoc "speedup" numbers. The paper proposes Time‑To‑Acceptance (TTA) under a documented acceptance test as the primary metric, together with a modeling framework that (a) reallocates task complexity across human/AI workflow stages, (b) makes rework explicit via acceptance probabilities, and (c) captures non‑linear scaling through effective parallelism and collaboration efficiency. To operationalize this, the authors introduce the Human‑AI Productivity Card (HAP Card): a lightweight, standardized reporting artifact to enable cumulative, comparable evidence during the pre‑AGI transition.
Key Points
- Current practice: reported speedups are often incomparable and misleading because they
- do not hold task complexity constant,
- trade quality for speed without explicit verification,
- omit rework, integration, verification, and governance costs that dominate deployment, and
- gloss over heterogeneity and uncertainty (task instances, workers, model stochasticity).
- Pre‑AGI context: general‑purpose AI is becoming cognitive infrastructure but is unevenly reliable, producing jagged, non‑monotonic performance frontiers across tasks and settings.
- Proposal: report productivity as Time‑To‑Acceptance (TTA):
- TTA is the expected wall‑clock time until a task output meets a pre‑specified acceptance test.
- Task complexity must be held invariant or explicitly normalized across comparisons.
- Heterogeneity (across tasks, workers, instances) and uncertainty (sampling, model stochasticity, nonstationarity) must be reported as first‑class quantities (distributions, confidence intervals).
- Rework explicitness: model rework with an acceptance probability p and explicit times for production, verification, and fixing; expected TTA increases with lower p and larger rework costs.
- Nonlinear scaling: introduce metrics like effective parallelism (how much parallel work actually reduces TTA) and collaboration efficiency (how well human and AI labor combine), since adding workers or models has diminishing and task‑dependent returns.
- Human‑AI Productivity Card (HAP Card): a standardized template listing task spec, acceptance test, time components, acceptance probabilities, uncertainty, model and prompt details, and costs to enable meta‑analysis and reproducibility.
Data & Methods (recommended framework)
- Primary metric
- Time‑To‑Acceptance (TTA): expected time from task start to an output that passes a documented acceptance test.
- Report mean/median and full distribution (or percentiles) and confidence intervals.
- Model of rework (simple representation)
- Let p = probability a candidate is accepted on an iteration.
- Let t_prod = time to produce a candidate, t_ver = verification time, t_fix = time to fix a rejected candidate.
- Expected TTA (simple form): E[TTA] = ((1 − p)/p) * (t_prod + t_ver + t_fix) + (t_prod + t_ver). (Derivable from geometric failures; report component breakdowns.)
- Report p and component times separately; sensitivity analyses when p depends on task instance or model version.
- Complexity normalization
- Hold task specification (input distribution, instructions, acceptance test) invariant across comparisons.
- When tasks differ, use explicit complexity normalization (e.g., information/content measures, standardized task buckets, cognitive work components) and report normalized metrics.
- Scaling and collaboration metrics
- Effective parallelism = realized speedup from adding parallel agents relative to ideal linear parallelism.
- Collaboration efficiency = ratio of achieved productivity to sum of isolated agent productivity (captures interaction overheads).
- Report how these metrics vary with team size, model scale/version, and task structure.
- Uncertainty & heterogeneity
- Report within‑task and across‑task variance, worker heterogeneity, model stochasticity, model drift (versioning), and sensitivity to prompts/data.
- Use randomized controlled designs where possible; pre‑register acceptance tests and analysis plans.
- Costs & externalities
- Report verification, rework, engineering/integration, monitoring, governance and compliance costs separately from raw generation time.
- Convert to economic measures when relevant (e.g., labor hours per accepted unit, $ cost per accepted unit).
- HAP Card (recommended fields)
- Task name and exact specification (dataset, input distribution)
- Acceptance test: exact rubric/automation used
- Model(s) & version, prompt/chain-of-thought/template used
- TTA: mean/median, percentiles, CI
- Component times: t_prod, t_ver, t_fix (distributions)
- Acceptance probability p (with CI)
- Effective parallelism and collaboration efficiency (with measurement protocol)
- Sample sizes, worker population, data splits
- Costs: time→$ breakdown (integration, monitoring, governance)
- Sensitivity analyses & caveats
- Links to data, code, and pre‑registered protocol
- Empirical design suggestions
- Compare human‑only, AI‑only, and hybrid workflows on identical task specs.
- Randomize across task instances and workers; report distributional effects (not just averages).
- Report model versioning and repeat key tests over time to capture nonstationarity.
Implications for AI Economics
- Measurement quality and growth accounting
- Standardized TTA and HAP Cards will improve estimates of productivity gains attributable to AI and avoid over‑estimating speedups that ignore quality/rework costs.
- Better microdata for growth accounting: more accurate contributions of AI to labor productivity and total factor productivity.
- Labor market modeling
- Explicit modeling of acceptance probabilities and rework clarifies which tasks are automatable and which require persistent human oversight, improving substitution vs. complementarity estimates.
- Heterogeneity reporting allows richer models of occupational impacts and wage dynamics by capturing variation in task complexity and worker skills.
- Investment and adoption forecasting
- Decision‑makers and investors will get more realistic expected time‑to‑value numbers, including integration and governance overheads, reducing misallocation driven by optimistic, non‑comparable speedup claims.
- Diffusion and scaling
- Metrics for effective parallelism and collaboration efficiency inform models of how productivity scales with deployment size and team composition, affecting industry structure forecasts.
- Policy and regulation
- Standard reporting helps procurement, safety, and regulatory oversight by tying claims to an explicit acceptance test and documented uncertainty; enables audits and comparability across vendors.
- Research & meta‑analysis
- HAP Cards create a corpus of comparable measurements to enable meta‑analyses and evidence accumulation during the pre‑AGI transition, improving the scientific basis for economic modeling and policy.
Summary: Treat human‑AI productivity as a measured, uncertainty‑aware scientific claim. Use TTA under a documented acceptance test, explicitly model rework and scaling, normalize for task complexity, and adopt HAP Cards to make cross‑study comparisons and economic inferences robust.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Human-AI productivity claims should be treated as scientific claims requiring standardized, complexity-normalized, uncertainty-aware reporting rather than ad hoc "speedup" numbers. Research Productivity | positive | quality of reporting on human-AI productivity claims (proposed change in methodology) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| In the pre-AGI transition, general-purpose AI is becoming cognitive infrastructure across tasks. Automation Exposure | positive | extent to which general-purpose AI functions as underlying cognitive support across tasks |
Reading fidelity
high
Study strength
low
|
not reported
|
| General-purpose AI remains unevenly reliable, producing jagged performance frontiers and non-comparable results across studies. Output Quality | negative | reliability and comparability of AI performance across tasks/studies |
Reading fidelity
high
Study strength
low
|
not reported
|
| Reported gains can silently trade quality for speed and often omit verification, rework, integration, and governance costs that dominate real deployments. Output Quality | negative | tradeoff between speed (reported gains) and output quality plus omitted deployment costs |
Reading fidelity
high
Study strength
low
|
not reported
|
| Productivity should be reported as Time-To-Acceptance (TTA) under a documented acceptance test, with task complexity held invariant (or explicitly normalized) and heterogeneity and uncertainty treated as first-class quantities. Task Completion Time | positive | time until an output meets a documented acceptance test (Time-To-Acceptance) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper proposes a quantitative framework that reallocates (rather than erases) complexity across human/AI workflow stages. Task Allocation | neutral | distribution of task complexity across human and AI workflow stages |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The framework makes rework explicit via acceptance probabilities. Task Completion Time | neutral | probability of initial output acceptance and consequent rework requirements |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The framework captures non-linear scaling through effective parallelism and collaboration efficiency. Team Performance | neutral | how productivity scales non-linearly with parallelism and collaboration efficiency |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We introduce the Human-AI Productivity Card (HAP Card), a lightweight reporting artifact to support cumulative and comparable evidence on human-AI collaboration in the pre-AGI transition. Research Productivity | positive | standardized reporting of human-AI productivity studies (HAP Card adoption and information captured) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Adopting TTA, complexity normalization, explicit heterogeneity/uncertainty reporting, and the HAP Card will enable cumulative and comparable evidence on human-AI collaboration during the pre-AGI transition. Adoption Rate | positive | ability to accumulate comparable evidence across human-AI studies |
Reading fidelity
high
Study strength
speculative
|
not reported
|