Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We conduct an empirical study of 86,156 test-file patches from 33,596 agent-authored PRs across 2,807 GitHub repositories produced by five coding agents: OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code.
Dataset collection and descriptive statistics reported by the study (counts of patches, PRs, repositories, and agents analyzed).
The production technology is multiplicative and cognitive capital functions as collateral that determines the return to AI adoption.
Analytic model specification and derivation inside the paper (formal mechanism). No empirical data.
The model features two state variables per agent, cognitive capital and cognitive debt.
Formal theoretical model presented in the paper (model specification). No empirical sample; analytic construction.
A pilot deployment in Newham's secure environment evaluated operational performance relative to manual workflows.
Paper reports a pilot deployment and an operational evaluation comparing DOMUS to existing manual workflows (method: pilot deployment; specifics such as duration or sample size not stated in provided text).
The authors evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels on LabOSBench.
Paper reports an evaluation suite covering multiple model classes and evaluation granularities (subtask and end-to-end); specific model identities and counts are not provided in the excerpt.
LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection.
Explicit specification in the paper of benchmark breadth: 96 subtasks and 8 simulators including the listed workflow stages.
Scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment.
Conceptual claim in the paper describing the nature of scientific-instrument operation; not supported by empirical data in the excerpt.
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems.
Statement in paper framing the problem; based on literature/field observation rather than reported experiments or data in this excerpt.
In evaluation runs, the evaluated model controls one coffee roaster while the remaining firms are controlled by fixed reference agents.
Experimental setup described in the paper: one roaster is controlled by the model under test; other five firms use fixed reference agents.
Each firm in CoffeeBench seeks to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing.
Specification of agent objectives and state variables in the benchmark design (cumulative net income objective; resources: cash, inventory; decision variables: pricing and transactions).
CoffeeBench simulates an economy of two farmers, two roasters, and two retailers operating autonomously over a 90-day simulation.
Environment description in the paper specifying the number and types of firms and the 90-day simulation horizon.
Frontier costs have stayed relatively stable between 2024 and 2026.
Authors' reported trend analysis of frontier model costs across the 2024–2026 period; excerpt lacks numeric cost trend data and sample size.
Every output is valuable (in H), trivial (in F \ H), or a hallucination (not in F).
Definition/assumption in the paper's formal model classifying outputs into three mutually exclusive categories relative to the formal language F and the valuable language H.
The empirical analysis used archival microdata from 770 large Spanish firms and employed staged OLS regression models.
Statement of data source and method in the paper's abstract.
The complementarity between AI deployment depth and breadth offers a configurational explanation for the AI productivity paradox.
Theoretical interpretation plus empirical finding of a positive interaction between depth and breadth in staged OLS analyses of archival microdata from 770 large Spanish firms.
AI capability can be conceptualized as two-dimensional: AI deployment depth (technological variety of AI implementations) and AI deployment breadth (organizational scope of AI diffusion).
Theoretical framing drawing on Resource-Based Theory and organizational search theory; conceptual argument presented in the paper.
Those preliminary experiments do not establish behavior preservation, scaling economics, or verified-change cost.
Authors' explicit limitation statement following the preliminary QLoRA experiments.
The review focuses on three core dimensions of impact: employee attitudes (job satisfaction, motivation, adaptability), workplace behaviours (performance, creativity, technology adoption), and organisational dynamics (leadership, trust, team cohesion).
Stated scope and focus areas in the abstract describing the review's analytical framework.
The study contributes a structured framework that clarifies the role of emotional AI in organisational contexts and outlines actionable, scalable strategies for real-world application.
Authors claim to have developed and presented a framework and strategies as part of the review paper (a descriptive/conceptual contribution rather than empirical evidence).
The study identifies key patterns, methodological trends, and underexplored areas in research on emotional AI systems in organisational contexts.
Authors report findings of their comparative analysis of the state-of-the-art literature (exact patterns/trends and counts are detailed in the full review).
This study follows the PRISMA framework to conduct a systematic evaluation and comparative analysis of the state-of-the-art literature on emotional AI in organisations.
Statement of methods in the abstract that the review used PRISMA; implies structured search, screening, and selection procedures described in the full paper.
The literature on AI-powered emotional intelligence systems is fragmented and insufficiently synthesised.
Authors' assessment based on a systematic literature review conducted following the PRISMA framework (details of databases, search terms, and included studies reported in the paper); exact number of studies not stated in the abstract.
The paper derives tight conditions that determine whether the economy is partially versus fully automated in the long run.
Analytical characterization in the model: derivation of necessary and sufficient (tight) conditions separating long-run partial automation from full automation (mathematical proofs within the paper).
Data accumulates endogenously as a byproduct of economic activity.
Model assumption and mechanism in the theoretical dynamic model: data generation is modeled as an endogenous outcome of agents' economic activity (analytical model specification).
Data is heterogeneous and task-specific.
Model assumption stated in the paper's setup: the model is built with data that varies across tasks and is task-specific (analytical model specification).
Density-normalized outcomes (e.g., smells per LOC) can mislead when treatment affects system size; raw counts and explicit decomposition are required for causal mining studies of AI tool adoption.
Interpretation and methodological recommendation derived from the observed pattern (unchanged smell counts + increased LOC leading to lower density) in the paper's empirical results.
Per-type estimates and robustness checks (wild cluster bootstrap, Lee bounds, stale-observation sensitivity) corroborate the main pattern; pre-trends are flat (Wald p = 0.90), consistent with the parallel trends assumption.
Placebo and robustness analyses reported in the paper (per-type breakdowns and multiple sensitivity checks) applied to the 151-repository panel; pre-trend test result reported as Wald p = 0.90.
Total architectural smell counts are essentially unchanged after adoption (+1.1%, p = 0.82).
Estimated treatment effect from staggered DiD / Borusyak imputation on total smell counts using the 151-repository panel (74 treated, 77 controls).
The causal effect of adoption on architectural smell density (ASD) was estimated using a staggered difference-in-differences design and the Borusyak imputation estimator.
Methodological claim describing the causal identification and estimation strategy applied to the 151-repository panel.
We mined 151 open-source Java repositories, 74 with detectable agentic AI adoption (identified via configuration files and Co-Authored-By commit trailers) and 77 propensity-matched controls, across a 13-month per-repository window yielding 1,811 monthly Arcan snapshots.
Descriptive dataset and methods statement in paper: 151 repositories (74 treated, 77 matched controls), 13-month windows, producing 1,811 monthly Arcan snapshots.
Causal evidence on the effect of AI coding tool adoption on software architecture is scarce; prior causal work has focused on code-level outcomes (complexity, static analysis warnings) and whether such degradation propagates to architecture-level outcomes remains unknown.
Literature/background statement in abstract asserting gaps in prior work; not an empirical result from this paper's dataset.
The superior zero-shot forecasting accuracy of foundation models does not inherently translate into better decision utility for resource consolidation.
Empirical analysis in the paper mapping forecasting outputs to downstream consolidation decisions and utility metrics, showing lack of improvement in decision utility despite better forecast accuracy.
We analyze 15,549 agentic PRs from 148 projects in the AIDev dataset.
Descriptive statement of dataset and sample used in the study (paper reports analysis of 15,549 agentic PRs from 148 projects).
From 2024 to 2026, more than 130 articles were submitted to this Special Issue (SI), and only 18 papers were accepted after rigorous peer review.
Editorial report in the paper describing CFP submissions and acceptance counts.
We conduct a qualitative study on a representative sample of 306 non-merged pull requests created or co-authored by the agents mentioned earlier, followed by a quantitative analysis of the reasons for rejection.
Authors' reported methods: qualitative study of a sample of 306 non-merged PRs and subsequent quantitative analysis.
Professional radiologists analyzed chest X-rays with access to state-of-the-art machine learning predictions in this replication setting.
Description of the experimental/contextual setting in the paper: professional radiologists using ML predictions on chest X-rays drawn from Collab-CXR.
The replication uses radiologist assessments from repeated-case designs, which include 68 radiologists and 11,420 paired radiologist–patient–pathology observations.
Direct reporting of study sample and design in the paper (repeated-case design; counts of radiologists and paired observations).
This note leverages the public Collab-CXR data repository described by Moehring et al. (2025) and first analyzed for human-AI collaboration by Agarwal et al. (2023).
Explicit statement in the paper identifying the data source and prior analyses (references provided).
In a production switchback experiment, the offline-trained policy reduces courier-side time costs without degrading customer-facing delivery quality.
Empirical claim supported by production switchback experiment described in the paper; asserts no degradation in customer-facing delivery quality concurrent with courier-side time improvements (no numerical metrics or sample sizes provided in excerpt).
The study uses a qualitative, mixed-methods design combining a systematic literature review, secondary evidence from an industry MRO digital survey, five semi-structured expert interviews, and two technical case studies (neural networks for aircraft retirement and an AI-based digital twin for a Power Electronics Cooling System).
Methods description provided in the paper (explicit counts: 5 interviews, 2 case studies); method = author-reported study design.
After screening, 35 studies were included in the thematic synthesis and supplemented by official regulatory and industry documents.
Review screening result reported in the paper: number of included studies = 35; supplementation by regulatory and industry documents stated.
A structured search protocol was designed for Scopus, Web of Science, PubMed, IEEE Xplore, and Google Scholar covering January 2016 to May 2026, English-language records only.
Methods statement in the review describing the databases, date range, and language restriction used for the systematic search.
The implementation literature on AI for pharmacy inventory and pharmaceutical supply chains remains dispersed across pharmacy operations, operations research, health informatics, and supply chain analytics.
The review's thematic synthesis of the searched literature (review methods described below) identified studies across these disciplinary areas.
Devil's Advocate (DA) is an AI assistant that critiques the human's initial ideas, whereas Dialectical Inquiry (DI) provides alternatives and synthesizes a resolution.
Conceptual/definitional claim in the paper describing the operationalization of DA and DI for the experiments.
This research empirically compares DA and DI in AI contexts.
Paper reports experimental comparison between AI behaviors implementing Devil's Advocate (DA) and Dialectical Inquiry (DI) across the studies.
Both studies examine benefit (information elaboration) and cost (cognitive load) pathways when AI supports SDM.
Paper explicitly frames both studies to measure information elaboration as a benefit pathway and cognitive load as a cost pathway; stated measurement plan in methods.
Study 2 tests mind-shaping interventions through user strategy training.
Study design described in the paper: a second experiment (Study 2) manipulating user strategy training (mind-shaping) to evaluate effects on SDM processes and outcomes.
Study 1 tests tool-shaping interventions by comparing three AI bot prototype conditions (Information-only, DA, DI) against a control treatment.
Study design described in the paper: randomized/controlled experiment (Study 1) with four conditions (three AI prototype conditions plus control).
The 'do no harm' property is confirmed empirically.
Abstract states empirical confirmation in simulations and applications; specifics (e.g., datasets, sample sizes) not included in abstract.
Including AI predictions as covariates has a 'do no harm' property: the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative.
Stated theoretical property in the paper and described as empirically confirmed in simulations and applications (per abstract).