Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
With the instruction files, 27.7% of the projects increased their merge rate by at least 20%.
Reported proportion of projects showing an increase in merge rate after creating instruction files based on the pre/post comparison of projects in the dataset (148 projects, 15,549 PRs).
Zhu et al. (2026) through a longitudinal case study propose a four-phase staged model (Digital Element Sedimentation; Digital-Intelligent Formation; Convergent Network Integration; Smart Engine Leap) explaining how traditional printing firms undergo digital transformation.
Longitudinal single-case study of Century Innovation as described in the paper.
Li et al. (2026) identify three cross-level models for high growth among digital startups (resource network orchestration model, innovation resource development model, entrepreneurial spirit coherence model) and find digital resource integration capability is a universal condition underpinning entrepreneurial growth.
Empirical/theoretical cross-level analysis reported in the paper (methods and sample size not given in excerpt).
Ma et al. (2026) show via machine learning and text analytics that an innovation culture stimulates R&D investment and aligns firm strategy toward digitalization; this positive impact is amplified by government innovation attention, especially in tech-intensive and non-state-owned enterprises.
Empirical analysis using machine learning and text analytics as described (specific sample size not provided).
Zhao et al. (2026) using social network analysis find robust peer effects within common ownership networks: firms in central network positions are more susceptible to peer influence, and industry leaders create demonstration effects that accelerate digital transformation among followers.
Social network analysis reported in the paper (methods referenced; sample size not provided in the excerpt).
Xu and Li (2026) demonstrate that technological core executives accelerate industrial AI transformation by leveraging parent–subsidiary executive connections and mobilizing subsidiary resources, improving supply chain efficiency and digital innovation.
Sector-specific empirical study in manufacturing reported in the paper (methods referenced; sample size not provided in excerpt).
Xu et al. (2026) find a positive relationship between top management teams' technological orientation and digital transformation investment in Chinese state-owned enterprises, conditional on managerial myopia, organizational slack, and environmental uncertainty.
Empirical analysis of Chinese state-owned enterprises reported in the paper (method and sample size not specified in the excerpt).
Chou et al. (2026a, 2026b) show in internet-only banking that perceived trust and service quality are fundamental determinants of consumer intention, while social influence plays a context-sensitive role; results validated with K-means clustering and a machine learning-assisted analytical approach.
Empirical analysis using Stimulus-Organism-Response framework, SEM/ML methods, and K-means clustering as reported in the paper.
Chen et al. (2026) find that argument quality and source credibility significantly enhance subscription intentions and willingness to pay more for GenAI tools among cross-border e-commerce operators; internet celebrity and user endorsements are key drivers of credibility.
Empirical study using the Elaboration Likelihood Model as described in the paper (study specifics such as sample size not reported in the text).
AI can better emulate cognitive functions that were once the exclusive domain of human experts.
Literature citation (Huang and Rust, 2021) and summary statement in the paper.
AI adoption enables managers to more efficiently use vast data streams for innovation and strategic orientation.
Cited conceptual/empirical claim attributed to Luo and Wang (2026) and discussion in the paper summarizing literature and observed trends.
In 2025, China’s digital economy contributed nearly 45% of the national GDP.
Stated in paper as a factual statistic (cited as [1]); presumably based on national economic statistics for 2025 referenced by the authors.
There is a need for good prioritization of tasks so that generated fixes do not lead to wasted human review efforts or wasted agent resources (e.g., tokens, compute, or allowed number of requests).
Authors' conclusion based on observed rejection rates and categorized reasons; recommendation to reduce wasted review/agent resources.
The results suggest improving model guidance by (1) proposing hints about the approach to follow for fixing an issue, (2) outlining constraints or limitations regarding approaches that should not be taken, and (3) instructing the agent on how to validate the implementation through CI pipelines and without introducing a breaking change.
Recommendations derived from the qualitative findings and analysis of rejection reasons in the paper.
AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects.
Introductory statement in the paper; no quantitative backing provided in the excerpt.
Higher belief calibration predicts larger incremental value from AI assistance.
Analysis of the Collab-CXR repeated-case assessments using the same analytic framework as Caplin et al. (2025b); includes 68 radiologists and 11,420 paired observations.
Lower baseline ability predicts larger incremental value from AI assistance.
Empirical analysis of repeated-case radiologist assessments from the Collab-CXR repository following the methods of Caplin et al. (2025b); sample includes 68 radiologists and 11,420 paired observations.
The results of this replication support the external validity of Caplin et al. (2025b)'s core findings: lower baseline ability and higher calibration predict larger incremental value from AI.
Reproduction of the analysis approach in Caplin et al. (2025b) applied to the Collab-CXR repeated-case assessments; uses radiologist assessments from the repeated-case designs (68 radiologists, 11,420 paired radiologist–patient–pathology observations).
On a diagnostic synthetic dataset tailored for MAS (explicit task decomposition, context separation and parallelization potential), expert-architected MAS consistently outperforms automatically generated architectures in both raw performance and cost-efficiency.
Controlled experiments on a synthetic diagnostic dataset designed by the authors to expose MAS advantages, comparing expert-designed MAS vs automatically generated MAS (experimental results reported in the paper).
Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making.
Statement of background / literature consensus in the paper (no specific empirical study cited in the excerpt).
Results illustrate how world feedback from a live economic and logistics system can be used to safely adapt decision policies online.
Generalization drawn from the reported deployment and switchback experiment showing practical use of delayed operational outcomes as feedback to adapt policies; presented as an implication of the empirical results in the paper.
In a production switchback experiment, the offline-trained policy reduces courier-side time costs.
Empirical claim from the production switchback experiment in the paper; reported reduction in courier-side time costs but no numerical magnitude or sample size provided in the excerpt.
In a production switchback experiment, the offline-trained policy increases batching.
Empirical claim supported by a production switchback experiment reported in the paper; the excerpt states an increase in batching but does not provide sample size or numerical effect size.
We use Double Q-learning targets and a conservative regularizer to reduce out-of-distribution value overestimation.
Methodological claim: inclusion of Double Q-learning targets and a conservative regularizer in training to mitigate overestimation; no quantitative measurement of reduction provided in excerpt.
We train a shared value function using centralized offline data and decentralized store-level execution.
Methodological statement describing training procedure: centralized offline dataset used to train a shared value function that is executed in a decentralized fashion at store level (no performance numbers given in excerpt).
This interface enables offline policy learning under noisy, delayed, and coupled feedback while preserving production feasibility constraints and operational safeguards.
Claim based on system design: using the multiplier interface keeps combinatorial optimizer intact and allows offline learning from logged delayed signals while maintaining feasibility and safeguards (no quantified evaluation included in excerpt).
We present a deployed reinforcement learning system at DoorDash that adapts dispatch objective weights in a large-scale food-delivery marketplace using delayed signals.
Description of a production deployment and system architecture in the paper; stated as a deployed system adapting dispatch weights using delayed operational outcomes (no numerical sample size reported).
The same focal MRO's integrated quote-to-contract system generated $6M additional revenue.
Field implementation / technical case study in a focal MRO reported in the paper; method = case study (focal MRO). Single-firm implementation; sample size effectively 1 organisation.
A focal MRO's field implementation of an in-house planning platform produced $4.2M annual savings.
Field implementation / technical case study in a focal MRO reported in the paper; method = case study (focal MRO). Single-firm implementation; sample size effectively 1 organisation.
Predictive maintenance is the dominant digital priority (prioritised by 56% of respondents).
Industry MRO digital survey reported in the paper (percentage reported: 56%); method = secondary evidence from an industry MRO digital survey. Sample size not stated in abstract.
Our results highlight the importance of process-oriented evaluation for reliable assessment of multimodal engineering reasoning systems.
Conclusion drawn by the authors based on the presented dataset, framework, benchmarking results, and human-automated agreement metrics.
Human evaluation shows strong agreement with our automated framework, with a mean absolute error (MAE) of 0.67 on a 10-point grading scale.
Empirical comparison reported in the paper giving MAE between human graders and automated framework on a 10-point scale (MAE value provided; sample size not stated in the excerpt).
Human evaluation shows strong agreement with our automated framework, achieving a Pearson correlation of 0.975 with the automated scores.
Empirical comparison between human graders and the automated evaluation framework reported in the paper (Pearson correlation value provided; sample size of human evaluations not stated in the excerpt).
The 8-stage framework independently evaluates each stage of the solution, enabling fine-grained analysis of reasoning failures.
Methodological claim about the functionality of the evaluation framework as described in the paper (framework design and intended analytic granularity).
We introduce an 8-stage automatic evaluation framework for assessing VLM-generated solutions.
Methodological claim in the paper describing the design of an 8-stage automatic evaluation framework (paper method section).
EngVQA contains 696 problems.
Explicit dataset size reported in the paper ("containing 696 problems").
We introduce EngVQA, a multimodal benchmark for evaluating engineering reasoning across 5 engineering subjects.
Dataset/method claim in the paper describing the scope of the introduced benchmark (explicitly states '5 engineering subjects').
Engineering problem solving requires interpreting technical diagrams, selecting governing physical principles, and maintaining physically consistent multi-step reasoning.
Conceptual claim made in the paper describing the requirements of engineering problem-solving (methodological/definitional statement rather than empirical evidence).
Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks.
Statement in paper citing prior benchmark performance of VLMs on general multimodal reasoning tasks (background claim; no specific benchmark names or numeric results provided in the excerpt).
AI is most valuable when used to augment, rather than replace, human supply-chain judgment and when deployment follows an implementation logic centered on resilience, compliance, and measurable operational value.
Review conclusion drawn from the thematic synthesis of the 35 included studies and supporting regulatory/industry guidance.
Reported outcomes include better visibility for shortage planning.
Outcome summary across included studies and documents (reported in the review); specific metrics not given.
Reported outcomes include reduced manual stock tracking.
Outcome summary across included studies (reported in the review); likely from implementations automating inventory monitoring.
Reported outcomes include tighter control of perishability and lost sales.
Outcome summary across included studies (stated in the review); details on magnitude not provided here.
Reported outcomes include improved supply and inventory prediction accuracy.
Outcome summary across included studies (reported in the review); no aggregate accuracy metrics provided in the sentence.
Reported outcomes include lower average inventory costs.
Outcome summary across included studies (reported in the review); specific studies not quantified in the statement.
The field has moved from descriptive dashboards toward integrated architectures that combine machine-learning forecasting, mathematical optimization, simulation, reinforcement learning, and automated medication management.
Thematic synthesis of the 35 included studies and supporting regulatory/industry documents noting evolution of system architectures in the literature.
Artificial intelligence is increasingly proposed as a remedy for pharmacy inventory volatility, medicine shortages, and fragmented pharmaceutical supply chains.
Opening statement of the review synthesizing recent literature and proposals across pharmacy operations and supply-chain fields; no specific empirical study cited in the sentence.
In three wet-lab validation experiments, OpenAI's o4-mini-high produced scripts that, when run on an OpenTrons liquid handling robot, successfully assembled DNA with expected sequences.
Three wet-lab validation experiments reported in the paper where scripts generated by the o4-mini-high model were executed on an OpenTrons robot and yielded DNA assemblies matching expected sequences.
Agents performed highly on tasks drawing on published knowledge and well-documented protocols.
Reported breakdown of agent performance across task types in ABC-Bench showing strong results on protocol-driven / literature-grounded tasks (authors' summary statement in results).
All tested LLM agents outperformed the median expert human baseline on all three tasks.
Benchmark evaluation results reported in the paper comparing multiple LLM agents against a median expert human baseliner across the three ABC-Bench tasks. (The excerpt does not report number of agents or number of human experts.)