Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The findings are consolidated via the AI Engineering Integration Framework and the Skills Transition Risk Matrix, which provide guidelines for strategically harnessing AI while safeguarding the Engineering profession.
Paper reports development of two conceptual/practical tools (framework and matrix) as outputs of the study; no validation details provided in abstract.
Case studies were performed covering five major industries.
Paper's reported methodology (number of case studies stated in abstract).
A Delphi study was conducted with 40 global experts.
Paper's reported methodology (Delphi sample explicitly stated in abstract).
A comprehensive mixed-methods study was conducted, incorporating a survey of 320 organizations.
Paper's reported methodology (survey sample explicitly stated in abstract).
AwareLLM was evaluated in a user study with 20 participants, compared to a standard LLM assistant across multiple tasks.
Experimental methods statement in paper; explicitly reports a user study and sample size.
The dominant paradigm for AI agents is an "on-the-fly" loop in which agents synthesize plans and execute actions within seconds or minutes in response to user prompts.
Statement in paper presenting a characterization of current AI agent design; conceptual/observational claim with no empirical data or sample reported.
Once functional deployment and operational investment are controlled for, worker-task use is not associated with employment declines.
Multivariate regression results reported in the paper using BTOS AI supplement data showing the coefficient on worker-task use becomes statistically indistinguishable from zero after controlling for functional deployment and operational investment; exact model details and sample size not provided in excerpt.
The synthesis covers research and practitioner guidance from the years 2023–2025.
Methods statement specifying the temporal scope of sources used for the synthesis.
This paper synthesizes recent research and practitioner guidance (2023–2025) to develop a practical model for designing human–AI collaboration in the financial reporting function (controllership).
Methods section declaration describing scope and approach (literature/practitioner guidance synthesis covering 2023–2025).
We conducted a controlled experiment comparing traditional task-splitting methods with AI-assisted approaches using GitLab Duo.
Methodological statement in the paper reporting a controlled experiment using GitLab Duo; sample size not stated in the provided summary.
We validate the framework empirically on five benchmarks (MATH, MMLU, TriviaQA, SimpleQA, LiveCodeBench) across eight models from five providers.
Empirical experiments reported in the paper using five named datasets and eight models from five providers (experimental evaluation / benchmarking).
For k-model cascades, first-order conditions imply a single shadow price that equalizes marginal quality-per-cost across stage boundaries.
Analytical derivation of first-order conditions for k-stage cascades within the decision-theoretic constrained-optimization framework presented in the paper.
Given a pool of k models, the frontier achievable by deterministic two-model threshold cascades is the pointwise envelope over choose(k,2) pairwise cascades, with switching points where the optimal pair changes.
Theoretical characterization/derivation in the paper (mathematical result about deterministic two-model threshold cascades and combinatorial envelope over pairwise cascades).
Reciprocal shadow prices link the budget-constrained and quality-constrained formulations of the cascade optimization.
Analytical derivation in the decision-theoretic framework using constrained optimization and duality presented in the paper.
For a two-model cascade, the cost-quality frontier is piecewise concave on decreasing-benefit regions of the confidence support.
Theoretical development in a decision-theoretic framework using constrained optimization and duality; proven properties for the two-model case reported in the paper (analytical result).
Under standard smoothness and finite variance conditions, SGD is minimax optimal for finding stationary points measured by l2-norms, thereby fundamentally precluding any complexity gains for sign-based methods in standard settings.
Theoretical statement based on prior minimax optimality results for SGD under standard smoothness and finite-variance assumptions (as cited/used in the paper). No new experiment; relies on worst-case lower-bound theory.
The authors ran a within-subjects study comparing authoring AD from scratch against editing AI drafts of varying quality.
Explicit methodological statement in the paper (within-subjects study design); sample size not reported in the excerpt.
RefineAD is an editing interface for human revisions (used to compare human editing of AI drafts against authoring from scratch).
Description of the authors' interface/tool used in the study (methodological claim).
GenAD is an AD generation pipeline that incorporates accessibility guidelines and contextual video information.
Description of the authors' system design (methodological claim from the paper).
The analysis identifies three major thematic areas: integration of AI in global supply chains; challenges and opportunities associated with AI adoption; and the impact of AI on decision-making and operational efficiency.
Structured synthesis of themes across 31 scholarly sources included in the qualitative literature review.
The study uses panel data of A-share listed energy-intensive firms from 2009 to 2021; measures corporate digital technology integration by counting frequency of digital-technology-related words in annual reports (text analysis); and evaluates low-carbon transformation using the LTFP method.
Methods and data description provided in the paper's abstract/summary: panel of A-share listed firms in energy-intensive industries (2009–2021); text analysis of annual reports for digital technology integration; LTFP method for low-carbon transformation measurement.
The study tested Olava Extract against five frontier models.
Method statement in the paper/abstract specifying comparison with five frontier models.
Methodologically, the work integrates dual eye-tracking, pupillometry, episode-based analysis, and causal inference to capture SSRL as a dynamic, emergent process.
Description of methods and measurement approach across studies: dual eye-tracking, pupillometry (for JME), episode-based analysis, and causal modeling are reported as combined methodology.
The paper reports three eye-tracking studies involving 182 dyads engaged in collaborative debugging tasks.
Stated description of methods in the paper: three eye-tracking studies, total sample of 182 dyads, task = collaborative debugging.
Data analysis utilized regression modeling for performance correlations, time-series analysis for predictive maintenance patterns, and thematic analysis for qualitative interviews.
Paper methods: explicit listing of analytic techniques used (regression, time-series, thematic analysis).
Secondary data encompasses sustainability reports, carbon footprint assessments, and operational performance metrics.
Paper methods: explicit listing of secondary data sources (sustainability reports, carbon footprint assessments, operational metrics).
Blockchain transaction records spanning eighteen months across Nigeria were used as primary data.
Paper methods: explicit statement about 18 months of blockchain transaction records across Nigeria.
The study uses IoT sensor data from forty-five facilities.
Paper methods: explicit statement that IoT sensor data were collected from 45 facilities.
Primary data collection includes structured interviews with supply chain managers.
Paper methods section: primary data described as including structured interviews with supply chain managers (number of interviewees not specified).
The study uses mixed methods involving case studies from twelve multinational companies across the manufacturing, logistics, and retail sectors.
Paper statement of methods: explicit mention of mixed methods and case studies from 12 multinational companies across the three sectors.
In educational settings, the use of GenAI does not consistently translate into improved learning or skill development, highlighting the need for careful integration of GenAI into computer science education.
Meta-analytic finding of a non-significant pooled effect on learning (g = 0.14, 95% CI [-0.18, 0.47]) combined with interpretation and recommendation in the paper's discussion.
Risk of bias in included studies was assessed using RoB2 and ROBINS-I.
Methods statement in the paper indicating the use of RoB2 (for randomized studies) and ROBINS-I (for non-randomized studies) to evaluate risk of bias.
Studies were required to compare GenAI-assisted with unassisted programming using quantitative measures of productivity (task completion time, commits, lines of code) and learning (exam performance).
Inclusion criteria reported in the Methods section of the paper specifying the required comparators and quantitative outcomes.
GenAI assistance has no statistically significant effect on learning outcomes (Hedges' g = 0.14, 95% CI [-0.18, 0.47]).
Meta-analysis pooling studies that reported learning outcomes (exam performance) and computing a pooled Hedges' g with 95% CI; reported estimate crosses zero.
We conducted a meta-analysis of n = 23 studies reporting k = 27 effect sizes on GenAI-powered coding assistants.
Systematic literature search across ACM, arXiv, Scopus, and Web of Science for studies published between 2019 and 2025; inclusion criteria required comparison of GenAI-assisted vs unassisted programming and quantitative outcomes. The paper reports n=23 studies and k=27 effect sizes.
The paper constructs a firm-level measure of AI development using AI-related patent data from Chinese listed firms.
Descriptive/method section: AI-related patent data from Chinese listed firms used to construct a firm-level AI development measure.
The analysis uses over 23 million WIOA participation records from 2017–2023.
Statement in the paper about the data coverage: administrative records of WIOA participants totaling >23 million records across 2017–2023.
The paper introduces the 'Retrainability Index' to measure program outcomes using post-intervention wage recovery and shifts in Routine Task Intensity (RTI).
Methodological contribution described in the paper: formulation of a composite index (Retrainability Index) combining wage recovery and occupation RTI change to evaluate WIOA outcomes.
Technologically advanced firms operating in hypercompetitive markets gain little from AI adoption, reflecting diminishing returns from capability saturation.
Cluster-specific results from the multidimensional heterogeneity analysis indicating small or negligible TFP effects for clusters identified as technologically advanced and highly competitive.
The study employs a System GMM estimator to address potential endogeneity and uses Fixed Effects (FE) and Random Effects (RE) models for robustness checks.
Methodological statement in the paper describing the econometric approach; verifiable from the methods section (no sample size or instrumentation details provided in the supplied text).
Prompt-driven generation (even with detailed prompting) fails to address the central problem of architectural complexity management in AI-based software engineering.
Results showing prompting did not prevent code bloat/coupling; conceptual argument reframing the problem toward architecture management rather than prompt engineering.
Neither functional correctness nor detailed prompting mitigates this architectural decay in AI-generated code.
Experimental comparisons reported in the paper where functionally correct outputs and variants produced with more detailed prompting were evaluated for structural quality and showed persistent architectural degradation.
Existing literature has extensively examined general AI adoption but limited empirical evidence exists on how more autonomous, agent-like systems contribute to economic outcomes.
Literature review / positioning statement in the introduction of the paper.
The study uses panel data from the World Bank (World Development Indicators and Enterprise Surveys) and OECD AI indicators for the period 2015 to 2024.
Explicit statement of data sources and time period in the paper's methods section.
An AI Adoption Index was constructed using indicators of AI investment, business adoption, and innovation output as a proxy for diffusion of advanced AI capabilities (including agentic features).
Methodological description in the paper: index synthesis from OECD AI indicators and other measures of investment/adoption/innovation; exact index components and weighting described in methods (sample size not applicable).
User-defined constraint types maintain usability.
User studies report that despite the additional constraint-typing features, usability remained acceptable (details, metrics, and sample sizes not provided in excerpt).
We conducted a technical evaluation and user studies with general and expert participants.
Paper reports carrying out both a technical evaluation and user studies (methods section). Specific sample sizes not provided in excerpt.
Perceived usability and satisfaction among participants showed little difference across model sizes.
Reported participant-reported measures (usability and satisfaction) compared across model sizes 3B, 8B, and 70B for N=112 participants; paper states little difference across sizes (no numeric statistics provided in the excerpt).
We examine the performance of humans (N=112) assisted by RAG-assistants compared to LLM-only or LLM+RAG baselines.
Experimental comparison reported in the paper with N=112 human participants across conditions (human+RAG vs LLM-only vs LLM+RAG baseline conditions).
This work evaluates a chatbot-style assistant based on Retrieval-Augmented Generation (RAG) in a realistic multi-turn information-seeking scenario inspired by workplace settings where compliance with local legislation and secure handling of sensitive data are often key.
Reported experimental setup: a chatbot-style RAG assistant evaluated in a realistic multi-turn information-seeking scenario inspired by workplace settings (method description in the paper).