Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The paper uses qualitative content analysis, coding documents against the four analytical dimensions to generate a comparative typology of coordination approaches.
Method description: manual qualitative coding of the 36 documents into the specified dimensions, producing the typology distinguishing Chinese and U.S. approaches.
The study's empirical basis comprises 36 national‑level policy documents (18 from China; 18 from the United States) focused on scientific data governance.
Author‑reported dataset and sampling description in the Data & Methods section.
The comparative analysis is organized across four dimensions: coordination objectives, institutional actors, governance mechanisms, and stakeholder legitimacy.
Methodological design reported in the paper; documents were coded against these four analytic categories.
The authors recommend empirical approaches for future work including randomized controlled trials in labs, before-after adoption studies, and collection of microdata on instrument usage, model versions, and provenance to measure impacts.
Explicit methodological recommendations in the Measurement and empirical research agenda section; these are proposals rather than executed studies.
There is a need for rigorous evaluation metrics and benchmarks for safety, reproducibility, and empirical studies quantifying productivity or scientific impact of LLM-driven instrument control.
Identified research gaps and recommended empirical research agenda described by the authors; these are recommendations rather than empirical findings.
The evidence presented consists mainly of qualitative arguments drawn from documented advances and discussion of prototypes; no controlled experimental evaluation is presented.
Authors' own description in the Data & Methods section about the nature of evidence supporting their perspective.
This paper is a conceptual perspective/review rather than an original empirical study.
Explicit statement in the Data & Methods section that the contribution is a perspective synthesizing literature and illustrative examples with no controlled experimental evaluation.
Modern microscopes are increasingly software-driven and data-intensive, while existing ML tools for microscopy are task-specific and fragmented.
Synthesis of recent literature on optical microscopes, detectors, and task-specific ML for image analysis referenced in the perspective (descriptive claim; no new empirical data collected).
Techno‑economic assessments (TEA) and life‑cycle analyses (LCA) are necessary research tools to compare bio‑routes to incumbent chemical synthesis on cost and emissions, and current literature is incomplete in this regard.
Review notes the presence of some TEA/LCA studies but highlights gaps and heterogeneity in methods and results across case studies; many processes lack published TEA/LCA at commercial scales.
Analyses were conducted as intent-to-treat comparisons across arms, with hypothesis tests reported (including p-values) and principal stratification used for mechanism decomposition.
Methods statement: intent-to-treat comparisons, reported p-values for score differences, and use of principal stratification for separating total effect into adoption and effectiveness channels in the randomized trial (n = 164).
The primary outcomes analyzed were LLM adoption (use), exam score (grade points), and answer length.
Study’s stated primary outcomes in methods: adoption indicator, exam score on an issue-spotting exam, and answer length (measured). Sample size n = 164.
The study used a randomized controlled design with three arms: no LLM access, optional LLM access, and optional LLM access plus brief training.
Study methods description: randomized assignment of 164 law students to three experimental conditions as listed.
The intervention consisted of roughly a ten-minute training focused on how to use the LLM effectively.
Study description of the intervention in the randomized experiment (three-arm design with one arm receiving ~10-minute targeted training).
Findings are estimated for Chinese cities and require replication in other institutional contexts to assess external validity.
Scope statement in the paper — primary empirical sample limited to 274 Chinese cities; authors note generalizability limits and call for replication elsewhere.
The paper’s AI exposure index — capturing automation and service-sector transformation — is important for robust measurement in empirical work on AI’s macro and environmental effects.
Methodological claim justified by the paper's construction of the index and its use in the main and robustness regressions; robustness checks reported using alternative index specifications.
The paper constructs an AI exposure index that captures both industrial automation (robots) and AI-enabled transformation of service-sector jobs/tasks.
Methodological construction described in the paper combining measures of industrial robot adoption (sectoral push) and AI-driven changes in service-sector job/task content.
The study uses a panel of 274 Chinese cities from 2007–2021 as the primary empirical sample.
Descriptive dataset information reported in the paper — city-level panel covering 274 cities and the years 2007 through 2021.
Empirical validation of the book’s proposals would require complementary case studies, model documentation, and outcome measurements.
Author/reviewer recommendation in the blurb about methodological limitations and next steps; not an empirical finding.
The book is predominantly conceptual and policy-analytic and uses illustrative case vignettes rather than presenting a single empirical study.
Explicit methodological description in the Data & Methods blurb: synthesis of technical ideas, governance requirements, and illustrative vignettes; no empirical sample or experimental protocol described.
The evidence base is qualitative: the study uses conceptual framework synthesis, comparative analysis of multi-sector implementations, and case examples rather than randomized or large-sample empirical evaluation.
Methods and limitations section of the paper explicitly describing the evidence base and methods (qualitative synthesis, pattern extraction, cross-case lessons).
The paper presents a deployment pattern intended to be adapted by sector and regulatory context rather than a one-size-fits-all blueprint.
Explicit statement in the paper and the described pattern design; based on qualitative pattern extraction and prescriptive guidance.
Methodological claim: combining fixed-effects panel estimation, mediation analysis, and panel threshold models is an effective multi-method approach to (a) estimate average effects, (b) unpack causal channels, and (c) detect nonlinear stage-dependent impacts.
The paper's applied methodology: fixed-effects panel regressions, mediation framework, and panel threshold modeling on the 2012–2022 provincial panel.
The paper constructs a multidimensional digitalization index composed of digital infrastructure, digital service capacity, and the digital development environment.
Index construction described in data/methods: composite indicator combining measures of connectivity/broadband (infrastructure), e-commerce/digital finance (service capacity), and policy/institutional/human capital indicators (development environment).
The study is observational (panel) and subject to limitations: residual confounding is possible; two-way fixed-effects estimators can be biased with heterogeneous treatment timing or dynamics; external validity beyond China and non-grain crops is not established.
Authors' stated limitations and caveats in the paper regarding identification and generalizability of results from the CLDS 2014–2018 observational panel.
The study uses two-way fixed-effects (household and year) models as the primary identification strategy and employs propensity score matching (PSM) as a robustness check.
Methods section of the paper describing estimation strategy applied to the CLDS 2014–2018 panel of grain-producing households.
Attributing productivity changes specifically to AI requires causal identification beyond VIS accounting (e.g., experiments, instrumental variables, difference-in-differences).
Paper notes that VIS is an accounting framework and that causal attribution to AI requires econometric/experimental methods beyond input–output accounting.
The method uses BEA for industry output and industry-by-industry transactions, BLS for employment and hours worked, and IMPLAN for detailed input–output structure and sector mapping; coverage period is 2014–2023.
Explicit data sources and time coverage stated: public BEA, BLS, and IMPLAN annual data 2014–2023 used to construct input–output matrices and labor measures.
Limitations of the review include the small sample of studies, uneven geographic coverage, heterogeneity in methods across studies, and limited long‑run evidence (especially on generative AI), which complicate causal aggregation.
Author-reported limitations based on the meta-assessment of the 17 included studies (variation in methods, contexts, and time horizons).
Design of this work: a systematic literature review and meta‑synthesis of empirical findings from peer‑reviewed journals (2020–2025), based on 17 publications.
Stated methods and inclusion criteria of the paper: systematic review of peer‑reviewed literature (sample = 17).
Long-term evidence on generative AI’s structural labor‑market effects is scarce; few longitudinal studies exist.
Assessment of study horizons and methods among the 17 papers indicates limited long-run and longitudinal analyses specifically on generative AI impacts.
Empirical coverage is limited for low‑income countries; evidence from such settings is scarce.
Geographic distribution of the 17 reviewed studies shows concentration in advanced economies with few or no studies focused on low-income countries.
The literature shows a surge in research activity on AI and labor markets in 2023–2025 and a concentration of studies in advanced economies.
Meta-analytic summary of the publication years and geographic focus among the 17 selected publications (temporal and geographic count of included studies).
The paper proposes two conceptual models (AI/ML‑Driven Labor Market Transformation Model and Sectoral Impact and Resilience Model) to organize heterogeneous findings and generate testable hypotheses about how AI reshapes labor across sectors and skill levels.
Conceptual synthesis integrating Technological Determinism, Socio‑Technical Systems Theory (STS), and Skill‑Biased Technological Change (SBTC); the models are theoretical outputs of the review used to map mechanisms and heterogeneity rather than empirical findings.
There are substantial measurement and identification gaps in the literature: heterogeneity in measuring 'AI adoption', limited long‑run causal evidence, and geographic bias toward advanced economies.
Methodological assessment within the review noting variability across studies in AI measures (patents, investment, task exposure proxies), paucity of long‑run causal designs, and concentration of empirical studies in advanced economies; this is a meta‑evidence limitation statement.
The Iceberg Index indicates where capability exists but does not indicate whether or when job losses will occur.
Explicit caution in the paper noting the distinction between technical exposure (capability overlap) and realized labor-market outcomes; methodological limitation described.
The Iceberg Index captures capability overlap but does not capture firm adoption choices, regulatory constraints, social acceptance, complementarity effects, or worker reallocation dynamics.
Limitations section in the paper explicitly listing these omitted factors; methodological boundaries of the Iceberg Index stated.
Model and simulations are implemented with the AgentTorch framework.
Implementation note in the paper indicating AgentTorch was used to build the agent-based models and run simulations.
The simulation model represents 151 million U.S. workers as autonomous agents, covers 32,000+ distinct skills, links agents to thousands of AI tools, and provides county-level resolution (~3,000 U.S. counties).
Model specification described in the paper: large-population agent-based model (AgentTorch) parameterized with occupation, skills portfolios, wages, and county locations; counts provided in the paper.
The Iceberg Index is a skills-centered metric that measures the wage value of specific skills AI systems can perform within each occupation; it quantifies technical exposure (capability overlap), not displacement, adoption timelines, or realized outcomes.
Methodological definition: mapping of ~32,000 skills to occupations with wage-value contributions, summing wages of skills that current AI capabilities cover to compute the index.
Key empirical gaps remain: better measurement of K_T (AI/software capital), more granular matched employer‑employee and wealth data, and improved estimates of task-substitution elasticities are required to precisely quantify incidence and policy impacts.
Authors’ stated research agenda and limitations section, including sensitivity analyses showing outcome variation with parameter choices and measurement uncertainty.
The empirical strategy relies on fixed-effects regressions applied to a firm-level panel dataset spanning 2005–2024.
Methods statement indicating fixed-effects regressions on firm-level panel covering years 2005–2024.
The study operationalizes AI exposure using AI patent stock as a proxy and analyzes effects within task-based and skill-biased technological change frameworks.
Methods section: firm-level panel (2005–2024), AI patent stock used as proxy for exposure, analysis framed in task-based and skill-biased change theories, fixed-effects regression approach.
The analysis uses a 23-sector recursive dynamic CGE model calibrated to the 2019 Input-Output Table and simulated through 2035 for Vietnam.
Paper's methodological description of the model and calibration data (stated in the abstract/summary).
Brown AI is modelled as an exogenous investment surge in IT hardware and services in the CGE experiments.
Paper's scenario design and model specification describing S3.
Green AI is modelled as a Total Factor Productivity (TFP) shock applied to heavy manufacturing and electricity in the CGE experiments.
Paper's scenario design and model specification describing S2.
These findings carry implications for the integrated design of AI and biofuel support policies at both national and EU levels.
Policy conclusions in the paper derived from empirical results (elasticities, heterogeneity, mechanism analysis).
Estimation is performed using feasible generalized least squares (FGLS) with distributed lag specifications, controlling for Renewable Energy Directive shocks and common cyclical effects via time fixed effects.
Methods section: FGLS estimation with distributed lags; controls include Renewable Energy Directive shocks and time fixed effects.
The study uses a balanced panel dataset of 12 European Union countries over the 2008–2024 period.
Data description in the paper: balanced panel of 12 EU countries from 2008 to 2024 (inclusive).
The study employs a multidimensional clustering approach based on firm size, age, market competitiveness, and digital infrastructure to examine heterogeneous AI effects.
Methodological description in the paper stating multidimensional clustering variables (size, age, competitiveness, digital infrastructure) used to form firm clusters for heterogeneity analysis.
The study uses panel data of 3,366 Chinese A-share listed firms from 2015 to 2023.
Direct statement of dataset and time period in the paper (panel of 3366 Chinese A-share listed firms, 2015–2023).