The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

After ChatGPT's launch, U.S. firms that previously relied heavily on online contracted labor shifted spending toward AI providers and away from marketplaces; the most exposed firms raised AI spending by 0.8 percentage points and cut marketplace payments, implying modest but measurable substitution (about $0.03 of AI spending per $1 decline in contractor payments).

Payrolls to Prompts: Firm-Level Evidence on the Substitution of Labor for AI
Ryan Stevens · January 28, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Ryan Stevens unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Ryan T. Stevens provider ID
Using firm-level expense data and a DiD around ChatGPT's October 2022 release, the paper shows that firms with greater pre-shock reliance on online contracted labor increased spending on AI model providers and reduced marketplace labor spending, indicating partial substitution of outsourced tasks with generative AI.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Generative AI has the potential to transform how firms produce output. Yet, credible evidence on how AI is actually substituting for human labor remains limited. In this paper, we study firm-level substitution between contracted online labor and generative AI using payments data from a large U.S. expense management platform. We track quarterly spending from Q3 2021 to Q3 2025 on online labor marketplaces (such as Upwork and Fiverr) and leading AI model providers. To identify causal effects, we exploit the October 2022 release of ChatGPT as a common adoption shock and estimate a difference-in-differences model. We provide a novel measure of exposure based on the share of spending at online labor marketplaces prior to the shock. Firms with greater exposure to online labor adopt AI earlier and more intensively following the shock, while simultaneously reducing spending on contracted labor. By Q3 2025, firms in the highest exposure quartile increase their share of spending on AI model providers by 0.8 percentage points relative to the lowest exposure quartile, alongside significant declines in labor marketplace spending. Combining these responses yields a direct estimate of substitution: among the most exposed firms, a \$1 decline in online labor spending is associated with approximately \$0.03 of additional AI spending, implying order-of-magnitude cost savings from replacing outsourced tasks with AI services. These effects are heterogeneous across firms and emerge gradually over time. Taken together, our results provide the first direct, micro-level evidence that generative AI is being used as a partial substitute for human labor in production.

Summary

Main Finding

Firms that relied more on online labor marketplaces before ChatGPT (Oct 2022) materially substituted paid freelance labor with AI provider usage afterward. The substitution is heterogeneous: the most-exposed firms both ramped AI spending earlier and substituted labor at a far lower monetary cost — roughly $1 of online-marketplace labor replaced by only $0.03 of AI model-provider spending (highest-exposure quartile), implying very large implied cost savings. Aggregate patterns show online-labor-marketplace (OLM) spending share fell (0.66% → 0.14%) while AI model-provider share rose (0% → 2.85%) between Q4 2021 and Q3 2025.

Key Points

  • Natural experiment: introduction of ChatGPT (Oct 2022) used as a shock that increased AI awareness/adoption.
  • Firms with higher pre-ChatGPT shares of spending on online labor marketplaces (dosage) increased AI-provider spending more and reduced OLM spending more than less-exposed firms.
  • Magnitudes:
    • Highest-exposure quartile (≥75% OLM share in Q2 2022) increased AI-provider share by ~0.8 percentage points in Q3 2025 relative to least-exposed firms.
    • Highest-exposure firms cut OLM share by ~15 percentage points (absolute) relative to least-exposed firms (some evidence of pre-trend).
    • Substitution rate (δAI / δOLM): highest quartile ~0.03 (i.e., $1 labor ↓ → $0.03 AI ↑); middle quartile ~0.30. Estimates come from bootstrapped ratios (B = 500); uncertainty is sizable and heterogenous across quartiles.
  • Timing: higher-exposed firms ramp up AI spending earlier than lower-exposed firms.
  • Important caveat: paper observes only payments to marketplace freelancers and to identifiable AI providers (OpenAI, Anthropic). It does not observe many in-house AI costs (infrastructure, engineering headcount) nor spending on AI from other cloud/tech vendors that could be used for model access.

Data & Methods

  • Data source: Ramp expense-management platform; transaction-level data (corporate card + bill/ACH payments) matched to merchants via Ramp’s proprietary merchant database.
  • Sample window: Q3 2021 – Q3 2025.
  • Sample restrictions:
    • Include firms spending ≥ $2,500 in at least two consecutive months in 2021 (excludes late entrants).
    • For regressions, exclude firms spending < $2,500 in Q3 2025 (to limit exit bias).
    • Focus sample on firms that had positive OLM spending in both Q1 and Q2 2022 (pre-treatment period).
  • Merchant definitions:
    • Online labor marketplaces (OLM): Upwork, Fiverr, Toptal, PeoplePerHour, Arc, MarketerHire, Catalant.
    • AI model providers: OpenAI, Anthropic (excluded Google/Meta/Microsoft due to identification issues).
  • Outcome variables: firm-quarter shares of total spend devoted to OLM and to AI model providers.
  • Identification:
    • Difference-in-differences / Two-Way Fixed Effects (TWFE) design exploiting the timing of ChatGPT (Oct 2022) and cross-firm variation in pre-treatment OLM share as a continuous/bucketed “dosage” (quartiles).
    • Model: s^J_{i,t} = α_i + γ_t + Σ_k δ_k · I(k ≥ T_Q3_2022) · E_i + ε_{i,t}, estimated separately for J = OLM and J = AI; E_i = quartile of pre-period OLM share (Q2 2022).
    • Standard errors clustered at the firm level.
  • Substitution calculation: ratio of AI and OLM treatment coefficients (δ_AI / δ_OLM), bootstrapped (B = 500) to obtain confidence intervals.

Implications for AI Economics

  • Direct evidence of firm-level substitution: This paper offers rare transaction-level proof that firms are shifting discretionary spend from externally contracted labor (freelancers) toward pay-as-you-go AI model access after a salient generative-AI shock.
  • Large implied cost savings (at least for observable payments): The extremely low δ_AI / δ_OLM in the highest-exposure firms implies that observable AI-provider spending replaces a much larger dollar volume of freelance labor, consistent with large short-run cost-reduction potential from model usage (but see caveats on unobserved costs).
  • Heterogeneity matters: Exposure-based heterogeneity suggests substitution will be uneven across firms and tasks. Policy and labor-market forecasts should therefore account for compositional reallocation (who is exposed, and how firms internalize AI) rather than assuming uniform displacement.
  • Measurement implications: Studies that measure AI’s labor impact solely from postings or worker outcomes may miss the firm-side reallocation channels captured here. Expense-level datasets and careful merchant identification are valuable complements to occupation/task exposure indices.
  • Limits to inference about aggregate employment:
    • The observed micro-level substitution does not automatically imply net job losses—demand for AI-related engineering and in-house roles, new tasks, or increased demand for complementary services could offset freelance reductions.
    • Important unobserved margins (in-house AI infrastructure, cloud, engineering headcount) may materially change the implied cost arithmetic and substitution rate.
  • Policy and research priorities:
    • Track in-house AI investments, engineering hiring, and cloud/infrastructure spending to fully account for the resource costs of AI adoption.
    • Link firm spending shifts to employment and wage outcomes to assess net labor-market effects and distributional consequences (especially for freelancers and early-career workers).
    • Investigate mechanisms behind heterogeneity (returns to scale in AI deployment, organizational capabilities, complementarities) to inform retraining/transition policies.
    • Broaden merchant coverage (other model providers and cloud vendors) and longer follow-up to capture medium-run equilibrium adjustments.

Limitations to keep in mind: sample restricted to firms already using online marketplaces; only two AI vendors identified; potential pre-trend concerns for the most-exposed firms; omission of in-house and other AI-related costs; outcome is spending shares, not direct employment or task-level substitution.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The design leverages high-frequency, firm-level payments data and a plausibly exogenous, widely-shared shock (ChatGPT release), which gives credible causal leverage; however, threats remain from selection into the expense platform sample, measurement error in AI usage (payments to model providers likely understate total AI adoption, excluding in-house models and informal use), possible concurrent shocks, and limited external validity — together these reduce confidence in strong causal generalization. Methods Rigormedium — Uses a sensible DiD/event-study framework with a continuous exposure measure and tests for dynamics and heterogeneity, but the abstract does not report key robustness checks (parallel pre-trends tests, placebo shocks, controls for concurrent shocks, alternative exposure definitions, or bounded-instrument approaches). The approach is rigorous in intent but depends on unreported implementation details and robustness analyses. SampleQuarterly firm-level payments from a large U.S. expense-management platform covering Q3 2021–Q3 2025, with spending classified into online labor marketplaces (e.g., Upwork, Fiverr) and payments to leading AI model providers; sample consists of firms that use the platform (size and sector composition not specified in the abstract). Themeslabor_markets adoption productivity human_ai_collab IdentificationDifference-in-differences exploiting the October 2022 public release of ChatGPT as an exogenous adoption shock; treatment intensity is the firm's pre-shock share of spending on online labor marketplaces (exposure), with event‑study/heterogeneity analyses to track dynamic responses. GeneralizabilitySample limited to firms using one expense-management platform — may overrepresent tech-savvy or expense-managed firms, Measures capture paid API/provider spending only and may miss internal AI development, open-source/self-hosted models, or unpaid use (understates true AI adoption), Focuses on online contracted labor marketplaces (Upwork/Fiverr) — does not capture offline/agency contractors or internal staff substitution, U.S.-centric and post-ChatGPT period; effects may differ in other countries or earlier/later adoption waves, Monetary substitution does not directly measure changes in output, quality, or hours worked, limiting inference about productivity gains

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We study firm-level substitution between contracted online labor and generative AI using payments data from a large U.S. expense management platform. Other other use of payments data to measure spending on online labor and AI model providers
Reading fidelity high
Study strength high
not reported
0.8
We track quarterly spending from Q3 2021 to Q3 2025 on online labor marketplaces (such as Upwork and Fiverr) and leading AI model providers. Other other quarterly spending on online labor marketplaces and AI model providers
Reading fidelity high
Study strength high
not reported
0.8
To identify causal effects, we exploit the October 2022 release of ChatGPT as a common adoption shock and estimate a difference-in-differences model. Other other causal effect estimation of AI adoption on spending patterns using DiD
Reading fidelity high
Study strength medium
not reported
0.48
We provide a novel measure of exposure based on the share of spending at online labor marketplaces prior to the shock. Other other pre-shock share of spending at online labor marketplaces (exposure measure)
Reading fidelity high
Study strength high
not reported
0.8
Firms with greater exposure to online labor adopt AI earlier and more intensively following the shock, increasing their share of spending on AI model providers. Adoption Rate positive share of spending on AI model providers
Reading fidelity high
Study strength medium
not reported
0.48
By Q3 2025, firms in the highest exposure quartile increase their share of spending on AI model providers by 0.8 percentage points relative to the lowest exposure quartile. Adoption Rate positive change in share of spending on AI model providers (percentage points)
Reading fidelity high
Study strength medium
0.8 percentage points
0.48
Following the shock, these higher-exposure firms simultaneously reduce spending on contracted labor at online labor marketplaces (significant declines in labor marketplace spending). Task Allocation negative spending on online labor marketplaces (contracted labor)
Reading fidelity high
Study strength medium
not reported
0.48
Combining these responses yields a direct estimate of substitution: among the most exposed firms, a $1 decline in online labor spending is associated with approximately $0.03 of additional AI spending. Task Allocation mixed substitution ratio between online labor spending decline and AI spending increase
Reading fidelity high
Study strength medium
$0.03 of additional AI spending per $1 decline in online labor spending
0.48
This implies order-of-magnitude cost savings from replacing outsourced tasks with AI services. Firm Productivity positive cost savings from replacing outsourced tasks with AI services (interpretive conclusion)
Reading fidelity medium
Study strength speculative
not reported
0.05
These effects are heterogeneous across firms and emerge gradually over time. Adoption Rate mixed heterogeneity and timing of spending responses (AI adoption and labor spending decline)
Reading fidelity high
Study strength medium
not reported
0.48
Taken together, our results provide the first direct, micro-level evidence that generative AI is being used as a partial substitute for human labor in production. Task Allocation positive use of generative AI as a partial substitute for human labor
Reading fidelity high
Study strength speculative
not reported
0.08

Notes