Evidence (1335 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
No single capability-sourcing route dominates under all conditions; the headline results depend on the corresponding model mechanisms.
Mechanism-knockout experiments that switch acquisition friction, absorption, open-weight cost reductions, and substitution on or off.
The route by which incumbents enter generative AI depends on contractibility: firms tend to partner when the capability can be rented through an API and tend to absorb when the capability is too tacit to rent.
History-friendly agent-based simulation of incumbent sourcing choices across build, partner, acquire, absorb, and wait routes; stated as Proposition P1 and reported as a headline finding.
Movement toward digital work is pathway-dependent rather than uniform across career transitions.
The study constructs consecutive same-person job transitions and calculates changes in job-title digitalization scores between source and destination jobs, summarizing these by occupational groups.
AI substitutes for routine cognitive and manual tasks, shifting worker duties toward nonroutinized, interpersonal, and creative tasks.
Presented as a task-based displacement claim; no task-level dataset or estimates are supplied.
Digital-era leadership research increasingly frames leadership as distributed and technologically mediated, with AI systems serving as decision aids, communication intermediaries, or partial substitutes for leader tasks.
Synthesis of emerging e-leadership and digital leadership literatures.
The combined regulatory regimes may induce firms to relocate activities, partition product lines, or maintain dual compliance tracks, thereby affecting the location of data processing and AI development.
Analytical inference from the interaction of extraterritorial obligations, compliance costs, and regulatory differences; no firm-level relocation data are reported.
AI perspective is primarily associated with entrepreneurial outcomes among established women entrepreneurs rather than equally across both venture stages.
Multi-group PLS-SEM analysis of new and established women entrepreneurs; the abstract reports that AI perspective becomes significant primarily in established ventures.
In settings where strategic and operational authority are fused, the decision to delegate tasks to AI or retain human discretion is endogenous to the GM's locus of authority.
Conceptual implication applying the locational-assumption finding to AI task allocation; it is not directly tested with AI deployment data.
The shift-share decomposition indicates that exposed tasks lose ground mainly through changes in which occupations are posted, while the task mix within surviving occupations remains broadly flat.
Shift-share decomposition of posting-share and within-occupation task-composition changes.
On longer Hard-Noise sessions, Plan-and-Act improved Gemma 4 26B coverage but increased extraneous valid calls.
A 60-task Hard-Noise session was evaluated under a fixed run-scope configuration, comparing ReAct with Plan-and-Act.
Performance gaps between models widen as tasks require more state recovery and dependency management.
ECC was compared across Easy, Medium, Hard, History, and Cross task suites, which progressively include longer dependencies and cross-task state.
Automation changes the composition of labor demand by substituting for some tasks, with consequences for economic growth and wage inequality.
The model explicitly includes automation as one of three digitalization mechanisms and traces its effects on labor demand, growth, and distributional outcomes.
The suitability mapping uses task-level capability importance weights, so the resulting scores indicate how well an agent matches capabilities that matter most for a task rather than whether the agent clears a fixed capability threshold.
The paper explicitly distinguishes capability levels from task-importance weights and states that the mapping does not test threshold attainment.
Most LLM harnesses also tended to always escalate to AlphaFold, achieving high accuracy through exhaustive tool use at maximum cost.
Observed tool-invocation strategies from LLM harness experiments, evaluated using execution traces and structured routing summaries.
In the paper's central numerical parameterization, low- and medium-autonomy uses make augmentation uniquely optimal, a high-autonomy use lies in the coordination region near the risk-dominance boundary, and still greater autonomy makes automation dominant.
Numerical illustration calibrated using professional-services revenue-to-payroll ratios, local-employment multipliers, operating margins, and task-exposure estimates translated through an explicit realization rate.
In the vanishing-friction limit, human augmentation is selected when the static tipping point is below one-half, while automation is selected when the tipping point is above one-half.
Risk-dominance implication stated in the aggregate-shock analysis, conditional on the Burdzy et al. fast-revision assumptions.
The interval supporting both automation and augmentation paths widens when firms place more weight on the market that later revisers will create, and switching costs narrow the interval.
Closed-form Corollary 1 under the fixed-wage linear-payoff benchmark. Equation (32) gives the overlap width and equation (33) gives the condition for it to be positive.
With forward-looking firms and staggered opportunities to revise production plans, the same inherited employment structure can support either an automation cascade or an augmentation recovery, depending on firms' expectations about later adopters.
Proposition 3 derives two perfect-foresight paths using discounted integrals of the relative payoff G(x), with conditions for an all-automation path and an all-augmentation path. The overlap is nonempty under sufficiently small switching costs.
Under the paper's production and complementarity conditions, the economy can have both a high-employment human-augmented equilibrium and a low-employment automated equilibrium, with a unique unstable interior threshold separating them.
Proposition 1 analytically classifies equilibria when κ > 0 and condition (12) holds. In the coordination region GA < 0 < GH, both endpoint equilibria exist and the interior equilibrium x* is unique and unstable under myopic adjustment.
AI adoption is shifting banking work away from routine and repetitive tasks toward analytical reasoning, technology fluency, professional judgment, relationship management, and complex problem solving.
The paper's synthesis of organizational AI-implementation reviews and banking literature; no primary data were collected in this study.
Changes in the dial propagate to security rankings and downstream portfolio composition in an exploratory backtest.
Exploratory downstream evaluation linking dial-induced changes in investment stance to security rankings and portfolio construction.
AI reshapes work heterogeneously across occupations and sectors rather than uniformly destroying jobs.
Descriptive comparative analysis of occupational AI exposure across six sectors using secondary international datasets and indices; no causal econometric identification.
When two returned products are near-identical visually, the higher-positioned product wins approximately two-thirds of their pairwise contests, whereas the better-fitting product wins only slightly more often than chance.
Pairwise analyses of near-identical returned products compare the effects of display position and relative query fit on which product receives the click.
At the catalog level, a product's overall tendency to look unlike its neighbors does not predict whether it is chosen; the positive association emerges only when comparing the same product across occasions with different returned neighbors.
The analysis separates product-level means from occasion-specific deviations and uses product fixed effects to identify within-product variation in local visual distinctiveness.
The distinctiveness effect appears in probability units but disappears in a conditional-logit specification.
The paper estimates both a linear probability model and a conditional logit, motivated by a model in which visual distance affects product individuation rather than inherent utility.
The public-data model is more suitable for early-stage planning, risk flagging, and scenario analysis than for autonomous procurement decisions.
The model achieved moderate log-scale fit but retained a median absolute percent error of approximately 51.3% on the dollar scale.
Algorithmic management is a central economic force shaping labor markets on platforms, including through task allocation, rankings, surveillance, and pricing, and it materially affects workers' bargaining power and market outcomes.
Conceptual implication drawn from the review's synthesis of algorithmic management and platform labor relations; no quantitative causal estimate is reported.
Age-based bans and administrative rules impose different costs and compliance strategies from enforceable duties of care, influencing firms’ regulatory optimization and innovation paths.
Comparative economic analysis of blanket platform restrictions, reporting requirements, and enforceable duties, including their implications for user bases, advertising revenues, monitoring, and algorithmic-safety investment.
The paper examines platform labour participation in relation to three household-division outcomes: domestic-service expenditure, spousal labour supply, and labour-force exit due to household care duties.
The authors construct proxy dependent variables from the CSS2023 survey and estimate logistic and ordinary least squares regression models with demographic, socioeconomic, household, and geographic controls.
The automation-augmentation paradox means that AI may substitute for some tasks while increasing the importance of human framing, exception handling, creativity, and accountability.
Conceptual synthesis based on Raisch and Krakowski (2021) and the socio-technical discussion of AI-enabled task execution.
Worker organization generated by mass education could influence firms' automation-adoption paths and task allocation, affecting the distribution of productivity growth between labor and capital, wage bargaining, and incentives to substitute automation for labor.
Theoretical implication and proposed research agenda; no firm-level adoption, wage, productivity, or labor-share estimates are reported.
Human capital analytics may complement automation and AI through upskilling and improved labor deployment, while also substituting for some HR tasks through automation.
Conceptual implications for AI economics presented in the paper; this is a proposed analytical framing rather than an empirically estimated result.
Algorithmic management functions both as a labour-control mechanism and as a regulatory technology that platforms can reconfigure in response to, or to circumvent, regulation.
Conceptual contribution grounded in the qualitative case study of platform adaptation during implementation of Spain's Rider Law.
Platforms responded to the Rider Law by restructuring their organisations to shift legal risk while preserving operational control over couriers.
Interpretive analysis of interview evidence concerning platform counter-strategies during implementation of the law.
Store-item demand is heterogeneous and seasonal, so equal replenishment rules may undersupply high-velocity products and overstock slower-moving items.
Store-item mean-demand heat map and monthly seasonality profile from 500 daily store-item series covering 10 stores and 50 items.
Organizations facing digital disruption can respond strategically either by developing resources and capabilities internally or by acquiring them externally through mergers and acquisitions.
The claim is based on the dissertation's resource-based view and dynamic-capabilities framing, supported by citations to Wernerfelt (1984), Henfridsson et al. (2009), and Teece (2007). It is a theoretical proposition rather than a reported empirical estimate.
AI-enabled sustainability applications shift labor demand toward data, sustainability analytics, assurance, and roles combining sustainability expertise with AI skills, while potentially displacing routine sustainability-monitoring work.
Conceptual synthesis of labor and human-capital implications; the paper reports no worker-level sample, wage estimate, or displacement estimate.
Blockchain does not replace external auditors; instead, it shifts their work toward assuring systems, protocols, code, and judgment-dependent elements that cannot be encoded on-chain.
Conceptual and normative analysis distinguishing automated audit assertions from assertions requiring professional judgment.
Human–AI collaborative decision-making comprises multiple paradigms distinguished by the relative contributions of algorithmic and human reasoning, indicating that no universal collaboration model is appropriate for all decisions.
Systematic review by Li and Tian covering 627 publications.
AI substitutes for routine time costs but complements human cognitive attention, implying that AI adoption should rebalance human and AI roles rather than pursue only labor-saving automation.
Conceptual analysis of complementarity and substitution using production/time-cost and attention-based frameworks.
AI has a dual role: it can automate tasks to improve time efficiency while also mediating, prioritizing, and shaping the allocation and capture of attention.
Conceptual synthesis of automation, attention economics, and AI-enabled information infrastructures.
The learned reinforcement-learning policies exhibit state-dependent liquidity-allocation and rebalancing behavior that responds to mispricing, gas costs, uncertainty, inventory exposure, and risk preferences.
The paper trains PPO and PPO_narrow agents in a model-based concentrated-liquidity environment and evaluates them across volatility, gas-cost, and risk-aversion scenarios. Each policy is evaluated on 1,000 independent simulations; the experiment includes 72 scenarios and 10 training seeds per scenario.
The experience-driven scaffold improved execution-level performance but did not lead to meaningful strategy revision.
The scaffold combined an experiment journal, skill library, and evaluator agent; trajectory evidence showed repeated within-strategy modifications and few or no adoptions of strategy-level suggestions.
Agents' default training strategies differ systematically by agent rather than being primarily determined by the task.
Comparison of default strategies across seven benchmarks and four base models; Claude Code trajectories predominantly selected full-parameter SFT, while Codex CLI trajectories predominantly selected PEFT.
The value of assistance depends on both the assistant model and the task, so a single general-purpose model leaderboard may be inadequate for organizational model selection.
Observed divergence in rankings across usage modes and task domains, including task-specific reversals and variation in rank correlations.
The paper finds that stronger work-related orientation is associated with stronger observable traces of prior human direction, while the observable iterative response differs between the two modes of use.
The study tests hypotheses H1 and H2 using aggregate AEI activity profiles separately for 1P API and Claude.ai, with specified delegation and iterative coproduction constructed from multiple behavioral metrics.
Frontier-model leadership was fragmented by task: Opus 5 led frontend coding, Fable 5 led repository-level coding, and GPT-5.6 Sol led agentic terminal work.
The paper compares model scores across public frontend, repository-coding, terminal-agent, and professional-work benchmarks.
The evaluated results do not identify one universally optimal intervention policy because task-quality gains, response time, and human-attention demands must be balanced.
Discussion of the lightweight intervention policy and reported variation in task gains, attention efficiency, response quality, and latency across configurations.
Held-out text-and-metadata logistic routers improved over Baseline by 7.2–37.5 percentage points in all six evaluated settings but remained 18.5–28.9 points below the fixed-order oracle.
Routers were evaluated on a narrower six-setting subset using a stratified 70/15/15 split, development-set selection, and refitting on train plus development data.
The gpt-oss-120b router reduced under-escalation to 18.0% but increased over-escalation to 33.3%, whereas the Tier-majority policy under-escalated on 27.4% and over-escalated on 12.5% of test problems.
The paper compares directional routing errors against the realized fixed-order oracle on the primary held-out test set.