The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Bigger AI models can backfire: when workers overestimate AI, increasing model scale can reduce joint productivity and hurt profits; firms should focus on aligning perceptions and incentives, not just buying larger models.

The Scaling Paradox in Human-AI Collaboration
Anyan Qi, Mengxin Wang · August 01, 2026
arxiv theoretical n/a evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Anyan Qi unresolved corpus identity
  2. Mengxin Wang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Anyan Qi provider ID
  2. Mengxin Wang provider ID
An analytical model shows that AI scaling only improves joint human-AI performance when humans correctly perceive AI capability—over‑perception can create a 'scaling paradox' where larger AIs reduce system performance and firm profit, while under‑perception blunts but does not reverse gains.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isolation; instead, it often works alongside humans, raising the question of whether these gains persist in human-AI collaboration. In this work, we develop an analytical model to examine when the empirical scaling benefits of AI translate into improved human-AI joint system performance. We demonstrate that the performance of a human-AI system can scale positively as the AI scales up-provided that humans have an accurate perception of the AI's capabilities. Human misperception, however, can fundamentally alter this relationship: i) when humans over-perceive the AI's capabilities, a scaling paradox may arise, in which greater AI scale reduces overall system performance and amplifies firm-level profit losses, and (ii) when humans under-perceive the AI's capabilities, performance still improves with scale but at a substantially slower rate. We further show that firms can actively manage these distortions through operational policies such as cost internalization and perception alignment, whose effectiveness depends on the economics of AI deployment and the direction of human misperception. These findings suggest that organizations may benefit more from managing the human-AI interface than from simply investing in larger, more expensive AI systems. More broadly, our results suggest that AI scaling should be viewed not only as a technological challenge, but also as a behavioral and operational one, and caution against the view that larger AI systems will automatically lead to better operational outcomes. Whether AI scaling creates value ultimately depends on how increased AI capabilities shape human beliefs and collaborative efforts.

Summary

Main Finding

AI scaling (larger models / more compute) improves standalone model capability, but those technical gains do not automatically translate into better human–AI system performance. Whether scaling helps operational outcomes depends critically on how humans perceive AI capability. If humans over‑perceive the AI, increasing AI scale can reduce total system performance and firm profit — a "scaling paradox." If humans under‑perceive the AI, system performance still improves with scale but much more slowly. Firms can mitigate these distortions via operational policies (cost sharing and correcting perceptions), and in many cases improving the human–AI interface yields more value than further scaling the AI itself.

Key Points

  • Integration of scaling laws into an operations model: the paper embeds a standard empirical scaling-law relationship (AI capability rises with scale) into a capacity-allocation model of human–AI collaboration.
  • Human effort is endogenous and scarce: workers decide how much effort to allocate per project vs. how many projects to attempt; AI support changes that tradeoff.
  • Human misperception is the critical behavioral mechanism:
    • Over‑perception (humans think AI is better than it is): workers withdraw effort too aggressively → may produce a non‑monotone relationship between AI scale and joint output (the scaling paradox). Firm profits can fall as scale increases in this region.
    • Under‑perception (humans think AI is worse than it is): workers over‑invest effort per project → scaling still helps but with attenuated gains.
  • Asymmetry of distortion: over‑perception is more damaging for both system performance and firm profit than under‑perception. Under‑perception can sometimes partially correct firm–worker incentive misalignments.
  • Firm-level effects and interventions:
    • Firms bear deployment costs of higher-scale AI; worker choices need not align with firm profit even under correct perceptions.
    • Two operational levers:
      • Cost internalization (shift some AI cost to workers): can curb over‑use and induce more worker effort per project. Optimal degree depends on AI cost and scale — full cost transfer may be optimal only when AI is cheap/small-scale.
      • Perception alignment (inform/educate/correct beliefs): especially valuable under over‑perception to prevent excessive effort withdrawal. Under‑perception alignment may not always raise firm profit because higher worker effort can conflict with firm incentives to scale throughput.
  • Managerial implication: investing in correcting perceptions and designing incentives/operational rules can be more valuable than simply increasing AI scale.

Data & Methods

  • Approach: analytic / theoretical modeling (no novel empirical dataset).
  • Model primitives:
    • AI capability parameterized as a function of scale s via a scaling-law form (scale → improved success probability with factor α).
    • Sequential human–AI workflow: human initializes, AI attempts task, human reviews/corrects as needed.
    • Human has finite effort capacity and allocates effort across projects (effort per project trades off success probability vs. throughput).
    • Human perception modeled as a (potentially biased) belief about AI capability (over‑ or under‑perception).
    • Firm incurs per‑project AI deployment cost that increases with scale; firm profit aggregates successful project value minus deployment cost.
  • Analysis: closed‑form comparative statics and qualitative characterization of when joint performance and profit are monotone in scale versus when non‑monotone (scaling paradox) arises; evaluation of two policy levers (cost sharing and perception alignment) and their effects under different misperception regimes.
  • Assumptions/limits: stylized sequential collaboration, representative worker decision (no heterogeneity), static setting (no learning dynamics modeled), analytic emphasis rather than calibrated empirical estimates.

Implications for AI Economics

  • Rethink ROI of scaling: macro or firm-level predictions that extrapolate model-level scaling laws to productivity gains risk overestimating returns unless they account for human behavioral responses and incentive structures.
  • Behavioral frictions matter for capital allocation: firms should measure and manage worker beliefs and effort allocation when deciding whether to invest in larger models versus complementary organizational changes.
  • Asymmetric risk management: prioritize detecting and correcting over‑perception (automation bias) because its welfare and profit harms are larger and can reverse the benefits of scale.
  • Policy and contracting levers:
    • Incentive design (cost sharing, bonuses, monitoring) can reshape worker effort and throughput; optimal design depends on AI cost and scale.
    • Information interventions (training, transparent performance feedback, calibrated confidence displays) can align perceptions — especially important when over‑perception is present.
  • For economists modeling AI-driven productivity effects: incorporate human perception and endogenous effort allocation into growth/productivity models and empirical identification strategies (e.g., measure perceived vs. actual AI accuracy, track effort allocation, run randomized information or cost treatments).
  • Operational recommendation for firms and policymakers: pilot deployments should jointly experiment with scale, cost‑sharing, and perception alignment (A/B test information treatments and pricing of AI usage) to find combinations that convert technical capability into realized productivity.

Brief caveats: results are from a stylized analytic model. Empirical validation across tasks, worker populations, and dynamic learning/adaptation is needed; heterogeneity and longer‑run learning could moderate or amplify the effects identified.

Assessment

Paper Typetheoretical Evidence Strengthn/a — Paper is an analytical/theoretical model without primary empirical causal estimation; claims are derived from formal model logic and illustrative references to empirical studies rather than new causal evidence. Methods Rigormedium — The paper develops a parsimonious analytical model incorporating scaling laws, endogenous effort allocation, and behavioral misperception, and derives comparative statics and policy implications; rigor appears solid for a theoretical contribution but no robustness checks, formal calibration/estimation, or extensive empirical validation are provided in the supplied text. SampleNo empirical sample; analysis uses a formal analytical model of a firm-worker-AI project workflow with AI scale s and scaling factor α, endogenous worker effort allocation, and parameterized misperception (over‑ or under‑perception); illustrative references to prior empirical/experimental studies are used for motivation. Themeshuman_ai_collab org_design productivity adoption GeneralizabilityResults derive from a stylized analytical model with simplifying assumptions about project structure, cost forms, and effort allocation—real workflows may be more complex., Behavioral misperception is modeled in a reduced-form way; real-world belief formation, learning, and heterogeneity across workers and tasks are omitted., Model abstracts from multi-agent interactions, market competition, dynamic learning, and technology adoption frictions that affect firm-level outcomes., Scaling law is treated as a parametric primitive; different domains and tasks may exhibit different scaling relationships or thresholds., Empirical magnitudes and boundary conditions are not calibrated, limiting quantitative prediction for specific industries.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In the analytical model, human–AI system performance scales positively with AI scale when humans accurately perceive the AI's capabilities. Organizational Efficiency positive Overall performance of the joint human–AI production system
Reading fidelity high
Study strength medium
not reported
0.12
When humans over-perceive AI capability, increasing AI scale can reduce total human–AI system performance over a range of AI scale, creating a scaling paradox. Organizational Efficiency negative Total performance or reward output of the human–AI system as AI scale increases
Reading fidelity high
Study strength medium
not reported
0.12
When humans under-perceive AI capability, system performance continues to improve with AI scale, but at a substantially slower rate than under accurate perception. Organizational Efficiency positive Human–AI system performance as AI scale increases
Reading fidelity high
Study strength medium
not reported
0.12
Human misperception reduces total reward output relative to perfect collaboration, but over-perception and under-perception have asymmetric effects: over-perception can reverse the benefits of scaling, whereas under-perception preserves monotonic improvement at a slower rate. Organizational Efficiency mixed Total reward output and its response to AI scale
Reading fidelity high
Study strength medium
not reported
0.12
At the firm level, worker over-perception of AI capability widens the gap between worker-optimal and firm-optimal behavior and amplifies the negative effect of the scaling paradox on firm profit. Firm Productivity negative Firm profit and worker–firm incentive alignment
Reading fidelity high
Study strength medium
not reported
0.12
Under-perception can partially benefit the firm by inducing workers to exert more effort per project and undertake fewer projects, thereby reducing exposure to AI deployment costs. Firm Productivity positive Firm profit and exposure to AI deployment costs
Reading fidelity high
Study strength medium
not reported
0.12
The optimal degree of AI cost internalization depends on AI deployment cost and scale: shifting the full deployment cost to workers can increase firm profit when AI is inexpensive or used at small scale, while firms should absorb a larger share of the cost when deployment is more costly or AI scale is higher. Organizational Efficiency mixed Firm profit under alternative AI cost-sharing policies
Reading fidelity high
Study strength medium
not reported
0.12
Perception alignment is particularly valuable when workers over-perceive AI capability because it prevents them from withdrawing effort too aggressively; under under-perception, alignment does not necessarily increase firm profit. Organizational Efficiency mixed Firm profit and worker effort allocation under perception-alignment policies
Reading fidelity high
Study strength medium
not reported
0.12
A randomized field experiment reported in the paper found that experienced open-source developers completed tasks 19% more slowly when AI tools were available, despite expecting AI to make them more than 20% faster and believing after the experiment that AI had saved them time. Task Completion Time negative Developer task completion speed
Reading fidelity high
Study strength high
19% more slowly
0.2

Notes