Evidence (8807 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Ninety independent agent runs built the same application from one detailed specification, each scored on a fixed 14-criterion functional rubric (42 point maximum) and a visual quality review.
Experimental study described in the paper: 90 independent agent runs, single specification, evaluated on a 14-criterion rubric (42-point max) plus visual quality review.
Merge and revert rates held steady.
Analysis of merge rates and revert rates over the study period showing no meaningful change after adoption/mandate.
The analysis panel comprises 802 developers and 196,212 pull requests spanning January 2024–April 2026.
Data description in paper: longitudinal panel of developers and PRs with specified date range and counts.
The study relies on secondary evidence from the U.S. Census Bureau, U.S. Bureau of Labor Statistics, OECD, IMF, Stanford AI Index, McKinsey Global Institute, NBER, and recent experimental research published from 2020 onward.
Explicit methodological statement in the paper.
Wages of labor that is used only in final goods production and is not displaced by AI increase in line with overall GDP.
Analytical economic model result indicating proportional wage growth for final-goods-only labor relative to GDP. No empirical sample reported.
Traditional approaches had 75% accuracy in risk prediction (as reported in the paper).
Observational comparison reported in the study between AI modelling and traditional approaches; reported accuracy for traditional approaches provided but no sample size or methodological details in the supplied text.
The AI premium is not present for loadings on casual use or open-weight (open-source) model use.
Decomposition analysis showing null or weaker relation between AI premium and loadings on casual usage metrics and open-weight/open-source model consumption.
We construct a high-frequency AI Factor from growth in tokens, dollars, and users, estimate firm-level AI Betas from stock return comovement, and characterize the AI Premium.
Methodological claim based on constructing a factor (AI Factor) using metrics of tokens, dollars, and users; estimating firm-level betas via stock return comovement.
The analysis uses 380 trillion tokens of realized AI consumption across more than four hundred large language models from the licensed proprietary OpenRouter dataset covering approximately 2 percent of current global monthly AI token consumption.
Descriptive statement about the dataset used: OpenRouter licensed proprietary dataset; 380 trillion tokens; >400 LLMs; coverage ≈2% of global monthly AI token consumption.
New generation panel data methods were applied, taking into account cross-sectional dependence and heterogeneity across countries.
Methodological description in the paper indicating use of advanced panel techniques that account for cross-sectional dependence (common shocks/spillovers) and heterogeneity (institutional/structural differences).
This study examines the determinants of economic growth in the 27 countries with the highest GDP for the period 2008–2020.
Study sample and period explicitly stated in the paper: the 27 highest-GDP countries, years 2008–2020.
This study systematically reviewed 194 peer-reviewed articles published between 2011 and 2025.
Statement in the paper's abstract describing a systematic review of 194 peer-reviewed articles (2011–2025).
Comments detected as likely to be generated by LLMs remain relatively stable over time.
Detector-based proxy analysis on repository comments across 2021–2025.
We analyze more than 930,000 agent-authored pull requests.
Descriptive statement about the dataset used for the study: an analysis of >930,000 pull requests authored by autonomous coding agents.
Inequality-adjusted income measures are themselves not new.
Literature/contextual claim made by the authors (statement that prior measures exist; no new empirical evidence required).
The empirical analysis uses data on industrial robots and Chinese listed companies covering the years 2006–2019.
Data description reported in the paper: matched dataset of industrial robot measures and observations for Chinese listed firms spanning 2006 to 2019.
We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase across three dimensions: requirement satisfaction level, reasoning, and code localization.
Study design reported in the paper: recruitment of 49 programmers, 148 NFR assessments, use of GitHub Copilot and iTrust codebase, and three specified assessment dimensions.
Evaluating how well LLM-based dialogue systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of system outputs and the quality of multi-turn interaction.
Methodological argument presented by the authors to motivate multi-turn study design (conceptual / methodological claim).
Non-Functional Requirements (NFRs) are inherently vague, context-dependent, and involve many parts of a program, making them difficult to assess with single-turn correctness benchmarks.
Conceptual claim motivating the study, based on properties of NFRs discussed in the paper (no empirical measurement reported for this claim).
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness.
Positioning statement in the paper's introduction / literature overview (no new empirical data reported for this claim).
We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks.
Description of empirical study scope in the paper (listed number of models and benchmarks and task domains).
Using panel data for 30 Chinese provinces from 2016 to 2024, this study constructs a marginal cost-based indicator of agricultural pollution–carbon reduction synergy (APCRS).
Methodological description in the paper: panel data covering 30 Chinese provinces for 2016–2024 and construction of a marginal cost-based APCRS indicator.
A Clopper–Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router.
Application of the Clopper–Pearson (binomial) confidence interval to estimate a finite-sample bound on beta; methodological proposal demonstrated in the paper.
Average pairwise error correlation rho cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates.
Theoretical/statistical construction and argument in the paper showing different joint error distributions can have same marginals and pairwise correlations but different all-wrong (beta) probabilities.
The study uses panel data covering 30 Chinese provinces from 2010 to 2022 to analyze AI's impact on value chain upgrading.
Stated dataset description in the paper (30 provinces, 2010–2022); used for the econometric analyses reported.
The study contrasts usage across three populations: external personal-account users, external organizational-account users, and workers within OpenAI using an automated, privacy-protecting pipeline.
Methodological description in the paper stating the populations compared and the privacy-preserving data pipeline used.
The position taken in the paper is platform-agnostic.
Author statement in the paper explicitly describing the position as platform-agnostic.
This study used survey data from 426 AI-adopting Chinese manufacturing firms and analyzed hypothesized relationships using hierarchical regression to isolate moderating effects of supply chain integration.
Methodological statement reported in the paper.
The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization.
Paper description of benchmark design and experimental protocol specifying controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization.
We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills.
Description of benchmark dataset introduced in the paper; the paper reports the benchmark contains 382 tasks, covers six professional roles and 22 procedural skills (dataset construction / annotation process described in methods).
When restricted to multi-commit PRs, the Copilot within-repo effect dissolves to +4.8 percentage points (p = 0.59).
Subset analysis limited to multi-commit PRs for Copilot in the AIDev dataset; reported point estimate and p-value.
Within-repository controls eliminate Devin's co-authorship gap, reducing it from +33.5 percentage points to +1.6 pp (p = 0.73).
Within-repo controlled regression/analysis on Devin PRs in the AIDev dataset showing adjusted effect and p-value.
The argument that automation is leading to a general decline in employment opportunities is not supported by actual facts and trends; rather, it is a product of the pervasive influence of technological fetishism.
Author's evaluative conclusion based on the reviewed theoretical and empirical studies; the excerpt provides no specific datasets or statistical tests supporting this rebuttal.
A conceptual framework is developed showing how digital infrastructure and institutional support mediate sectoral transformation.
Paper presents a conceptual framework (theoretical/modeling component) derived from empirical findings and policy analysis; this is descriptive rather than a quantified empirical result.
We conclude by outlining implications for designing and evaluating human-AI teams as socio-technical systems and for prioritizing longitudinal and in-context studies that capture how teaming evolves over time.
Authors' conclusions and recommendations based on the systematic review and observed gaps in the literature (noted need for longitudinal, in-context studies).
Bibliometric patterns suggest a shift since 2020 from foundational demonstrations in controlled settings toward applied, higher-stakes contexts where trust dynamics, communication, and ethical accountability more directly shape adoption and sustained performance.
Bibliometric analysis of the 104 studies showing temporal trends (pre- vs post-2020) in research contexts and topics.
Across studies, performance was the most frequently examined aspect, followed by trust, explainability and transparency, decision-making, and team processes.
Synthesis and frequency coding of outcomes/measured constructs across the 104 included empirical studies.
Gaming and entertainment, aviation, military and defense operations, emergency response and public safety, and healthcare also represented substantial portions of the literature.
Domain breakdown from the systematic review of 104 empirical studies (frequency counts by domain reported in Results).
Cross-domain and interdisciplinary studies were the largest category, representing broad workplace or team-based investigations not tied to a single industry and instead focused on general collaboration issues such as communication, teamwork, coordination, and coworker interaction.
Categorization / coding of the 104 included empirical studies; frequency counts by study domain reported in review.
We conducted a PRISMA-guided systematic review with bibliometric analysis of 104 peer-reviewed empirical studies published between 2015 and 2025 and identified through Engineering Village, IEEE Xplore, PubMed, ScienceDirect, and Web of Science.
Methods reported in paper: PRISMA-guided systematic review and bibliometric analysis; explicit statement of 104 peer-reviewed empirical studies and databases searched (Engineering Village, IEEE Xplore, PubMed, ScienceDirect, Web of Science).
Order, entropy, information, and useful energy are task-dependent and system-relative concepts whose meanings depend on the objectives of the system.
Conceptual argument and discussion in the paper about the context-dependence of informational and energetic notions within the proposed framework; no empirical evidence provided.
The paper's contribution is theoretical: it reframes the AI productivity debate beyond automation anxiety by linking technological change, income distribution and effective demand in a single analytical framework.
Author-stated contribution and framing in the conceptual review (description of scope and aim).
The review proposes possible indicators for future empirical research, including the productivity–real labour income gap and an absorption tension indicator.
Paper's methodological/propositional content (explicitly proposes indicators for empirical work).
The paper defines the 'Distributional Absorption Threshold of AI-Induced Productivity' as the point beyond which productivity gains associated with AI are no longer accompanied by proportionate increases in broadly distributed real purchasing power and household consumption.
Textual/definitional content of the conceptual review (the paper introduces and defines this concept).
We study shared-workspace human-AI teams using the Collaborative Gym environment with DiscoveryBench tasks.
Methodological description in the paper stating the experimental environment and task suite used.
We ran 1,482 sessions in our experiments.
Statement in the paper reporting total experimental sessions (Collaborative Gym + DiscoveryBench).
Experimental results were validated through bootstrap confidence intervals, multiple-testing corrections, and subperiod stability analysis.
Methods statement in the paper indicating use of bootstrap confidence intervals, multiple-testing corrections, and subperiod stability analysis to validate backtest results.
The article presents a comprehensive framework spanning macro-level governance principles and micro-level interaction typologies, illustrated through case examples from telecommunications, retail, insurance, and consumer products sectors.
Descriptive/methodological claim in the paper indicating the presence of a framework and sectoral case examples (no quantitative sample sizes reported for cases).
The article draws on Deloitte's 2026 Global Human Capital Trends survey of over 3,000 business leaders across 15 countries.
Methodological/data source statement explicitly provided in the paper.
The robot is initialized with the selected episodic memory before a new collaboration episode begins (i.e., memory reuse occurs at episode start).
Design and procedure reported in paper: authors state that the robot is initialized with the selected memory prior to new episodes as part of the intervention tested in the experiment.