The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Advances in large language models speed professional work: each year of model progress reduces task time by about 8%, driven roughly 56% by added compute and 44% by algorithmic gains. The benefits concentrate in non-agentic analytical tasks, and if current scaling continues the authors estimate roughly a 20% boost to U.S. productivity over ten years—an extrapolation that rests on strong assumptions.

Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks
Merali, Ali · December 24, 2025 · arXiv (Cornell University)
openalex rct medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Merali, Ali provider ID

Semantic Scholar

Latest observation:

  1. Ali Merali provider ID
A preregistered randomized experiment with 500+ professionals finds each year of LLM progress cuts task time by about 8% (56% from increased compute, 44% from algorithmic advances), with larger gains on non-agentic analytical tasks than on agentic workflows, and model-scaling extrapolations implying roughly a 20% U.S. productivity uplift over the next decade.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper derives `Scaling Laws for Economic Impacts' -- empirical relationships between the training compute of Large Language Models (LLMs) and professional productivity. In a preregistered experiment, over 500 consultants, data analysts, and managers completed professional tasks using one of 13 LLMs. We find that each year of AI model progress reduced task time by 8%, with 56% of gains driven by increased compute and 44% by algorithmic progress. However, productivity gains were significantly larger for non-agentic analytical tasks compared to agentic workflows requiring tool use. These findings suggest continued model scaling could boost U.S. productivity by approximately 20% over the next decade.

Summary

Main Finding

Access to progressively more capable LLMs materially raises professional productivity: each year of frontier model progress reduced task completion time by ≈8% (p ≈ 0.04). Across a large preregistered RCT (N>500), AI assistance raised base earnings per minute (EPM) by 81.3% ($0.56/min), improved expert-assessed quality by ~0.34 SD (0.60 points on a 7‑pt scale), and increased total earnings per minute (TEPM, including performance bonuses) by 146% ($1.06/min, p < 0.001). A tenfold increase in training compute is associated with a 6.3% reduction in task time. Decomposing frontier progress, ~56% of gains are attributable to compute scaling and ~44% to calendar-time (algorithmic/data/architecture) improvements. Gains are much larger on non‑agentic analytical tasks than on agentic, tool‑using workflows. Projecting these elasticities in an aggregate growth framework yields an estimated ~20% U.S. productivity boost over the next decade under low inference-cost assumptions.

Key Points

  • Experimental design and scale

    • Large preregistered RCT with >500 screened professionals (management, data analysis, consulting).
    • Participants randomly assigned to control or one of 13 LLMs spanning multiple release dates and training‑compute scales.
    • High‑powered incentives: $15 base per task + $15 bonus for high‑quality submissions (grades ≥5/7). Expert graders assigned scores.
  • Average AI uplift (pooled over models)

    • EPM +$0.56/min (81.3% increase; p = 0.001).
    • Quality +0.60 points on 7‑pt scale (~0.34 SD; p < 0.001).
    • TEPM +$1.06/min (146% increase; p < 0.001).
    • Composition: increases in speed contribute ~52.6% of TEPM gains; quality improvements ~47.4%.
  • Scaling laws (economic impacts)

    • Calendar‑time scaling: each additional year of frontier model progress ≈ 8% reduction in task time.
    • Compute scaling: a 10× increase in training compute → ≈ 6.3% reduction in time.
    • Decomposition: ~56% of observed gains driven by compute scaling; ~44% driven by algorithmic/calendar‑time progress.
  • Heterogeneity by task type

    • Non‑agentic (analytical/writing) tasks: much larger gains — TEPM +$1.58/min (p < 0.001).
    • Agentic (multi‑step tool use, procedural workflows): smaller and statistically weaker gains — TEPM +$0.34/min (p = 0.46); difference across task types significant (p = 0.043).
    • Interpretation: scaling so far rapidly augments analytical cognition; procedural agency/tooled workflows lag.
  • Quality vs. capability scaling

    • Autonomous model outputs’ quality scales with compute.
    • Human‑assisted (human+LLM) output quality is flat across model generations — evidence of a “human cap” where users attain a satisficing quality threshold rather than exploiting maximal model capability.
  • Aggregate projection

    • Using estimated elasticities in an aggregate growth framework (following Acemoglu 2024), the paper projects ~20% U.S. productivity growth over the next decade, conditional on continued scaling and low marginal inference costs.

Data & Methods

  • Sample and recruitment

    • 500 participants recruited mainly via Prolific; stringent eligibility (≥$40k salary, ≥1 year experience) and competency screening (≈90% screen‑out).

    • Final sample: experienced professionals (79% ≥3 years experience; 47% ≥5 years).
  • Tasks and classification

    • One task per participant drawn from profession‑specific, realistic workflows (management, consulting, data analysis).
    • Tasks categorized as agentic (require multi‑step tool interactions) or non‑agentic (analytical/writing).
  • Treatments

    • Random assignment to control (no bot) or one of 13 LLMs varying in training compute and release date.
    • Custom website logged usage; monitored practice run ensured protocol compliance.
  • Outcomes

    • Time taken, base Earnings Per Minute (EPM), Total Earnings Per Minute (TEPM, includes bonus), and expert‑assessed grade (0–7).
    • Primary estimands: effects of any AI vs control; elasticities of outcomes to model release month and to log(training compute). Combined specs decompose compute vs calendar‑time effects.
  • Identification & analysis

    • Pre‑registered RCT (AEARCTR‑0013743); approved by Yale HRPP.
    • Regression controls included profession, task, demographics, abilities, AI familiarity, country dummies; robustness checks in appendices.
    • Statistical significance reported (p‑values); figures/tables in main text and Appendix A.

Implications for AI Economics

  • Empirical link from ML scaling to economic output

    • Provides direct, experimentally grounded elasticities linking model capability (compute and calendar‑time improvements) to human productivity — useful inputs for macro/sectoral forecasting and welfare analysis.
  • Heterogeneous automation potential

    • Analytical cognition appears highly susceptible to current scaling paradigms; procedural, tool‑driven tasks remain more resistant. Policy and firm strategies should distinguish between task types when forecasting displacement, re‑skilling needs, and investment in automation infrastructure (agentic toolchains, APIs, UI/UX).
  • Importance of human factors and complementarity

    • The “human cap” on quality suggests realized gains depend on how humans interact with models (training, prompts, incentives, UX). Investments in user training, interface design, and workflow integration may unlock further gains beyond raw model scaling.
  • Role of incentives and measurement

    • High‑powered piece‑rate style incentives amplified TEPM results; translating these experimental rates to typical salaried settings requires caution. Still, the study demonstrates that speed+quality improvements can compound substantially under performance‑linked compensation.
  • Policy and forecasting

    • A ~20% U.S. productivity uplift over a decade (conditional) is economically large and suggests rethinking near‑term growth scenarios and labor market policy (education, tax/transfer design, social insurance for transitions).
    • However, assumptions matter: inference costs, broader diffusion, regulatory constraints, and development of agentic/tooled models will affect realized aggregate impacts.
  • Evaluation methodology

    • Supports “centaur evaluations” (human+AI) as economically relevant benchmarks. Benchmarks that evaluate models in isolation risk misestimating real‑world productivity impacts.

Caveats & Limitations

  • Sample scope: three high‑skill professions and a screened, experienced online sample — not representative of the entire labor force.
  • Interface limitations: participants used standard chatbot interfaces with limited tool access; agentic task gains may be understated relative to systems with richer tool integration or autonomous agents.
  • Incentive design: large bonuses (doubling pay) may exaggerate TEPM effects relative to typical workplace compensation structures.
  • Proxies for algorithmic progress: model release date is an imperfect proxy for non‑compute improvements and may confound many contemporaneous changes.
  • External validity: scaling projections assume continued predictable scaling, low inference costs, and broad adoption — all uncertain.

If you want, I can (a) extract the main regression coefficients and p‑values into a single table, (b) produce a one‑page brief oriented to policymakers, or (c) simulate alternative aggregate productivity scenarios under different inference‑cost and adoption assumptions. Which would you prefer?

Assessment

Paper Typerct Evidence Strengthmedium — Strong internal validity for short-term task performance due to preregistered randomized assignment and objective time measures across many models, but limited external validity: lab-style tasks may not translate to on-the-job productivity, quality vs time trade-offs are not fully addressed, subgroup and long-run effects are uncertain, and the headline macro extrapolation (≈20% U.S. productivity gain) depends on scaling assumptions and aggregation choices. Methods Rigormedium — High rigor in experimental design (preregistration, randomization across 13 distinct models, reasonably large N), and an explicit decomposition of compute vs algorithmic contributions; however, rigor is tempered by likely reliance on assumptions to map model differences into 'years of progress', potential measurement limitations (time as sole/primary productivity proxy), limited information on learning or repeated-use effects, and extrapolation methods for economy-wide impacts that require strong assumptions. SampleA preregistered sample of slightly over 500 professionals (consultants, data analysts, and managers) who performed standardized professional tasks while using one of 13 large language models; primary outcomes were task completion time and task-type-specific performance (non-agentic analytical tasks vs agentic/tool-using workflows); model-level data included training compute and vintage/algorithmic generation. Themesproductivity human_ai_collab IdentificationPreregistered randomized experiment: over 500 professionals were randomly assigned to use one of 13 LLMs that vary in training compute and vintage while completing standardized professional tasks; causal effects are estimated by comparing task completion time across model assignments (with task and participant controls) and decomposing overall model-progress effects into components attributable to compute increases versus algorithmic/vintage improvements using model metadata and scaling relationships. GeneralizabilityLab-style, time-limited tasks may not reflect sustained on-the-job productivity or complex workflows., Sample limited to consultants, data analysts, and managers — excludes many occupations and firm types., Task-time measures may not capture work quality, downstream coordination costs, or managerial impacts., Agentic/tool-using workflows in the study may not represent future integrated agent systems., Extrapolation to national productivity relies on scaling assumptions and adoption dynamics that are uncertain., Geographic/sample demographics unspecified; results may not generalize across countries or firm sizes.

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Each year of AI model progress reduced task time by 8%. Task Completion Time positive task time (task completion time)
Reading fidelity high
Study strength medium
n=500
8% reduction in task time per year of model progress
0.6
56% of the observed productivity gains were driven by increased compute. Task Completion Time positive share of productivity gains (task time reduction) attributable to increased compute
Reading fidelity high
Study strength medium
n=500
56% of gains driven by increased compute
0.6
44% of the observed productivity gains were driven by algorithmic progress. Task Completion Time positive share of productivity gains (task time reduction) attributable to algorithmic progress
Reading fidelity high
Study strength medium
n=500
44% of gains driven by algorithmic progress
0.6
Productivity gains were significantly larger for non-agentic analytical tasks compared to agentic workflows requiring tool use. Task Completion Time positive task time / productivity gains by task type
Reading fidelity high
Study strength medium
n=500
0.6
The paper reports a preregistered experiment in which over 500 consultants, data analysts, and managers completed professional tasks using one of 13 LLMs. Other null_result study participation / experimental design
Reading fidelity high
Study strength high
n=500
1.0
Continued model scaling could boost U.S. productivity by approximately 20% over the next decade. Fiscal And Macroeconomic positive U.S. productivity (aggregate productivity)
Reading fidelity high
Study strength speculative
n=500
approximately 20% boost to U.S. productivity over the next decade
0.1

Notes