Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Coordinated reduction in working hours helps maintain aggregate demand.
The paper's synthesis of historical transitions and pilot programs and argument about distribution of productivity gains; no quantitative evidence or sample sizes provided in the summary.
Gradual, policy-led reduction in standard working hours can preserve employment.
Claim based on examination of historical work-time transitions, contemporary pilot programs, and cross-sector implementation strategies referenced in the paper; no specific studies or sample sizes cited in the summary.
Systematic quality auditing should be standard practice for complex agentic tasks.
Normative recommendation based on the authors' methodological and empirical findings that auditing revealed substantial benchmark issues affecting evaluation of agent capabilities.
Re-evaluating on ELT-Bench-Verified yields significant improvement attributable entirely to benchmark correction.
Re-evaluation of agent performance on the revised benchmark which the authors claim shows significant improvement and that this improvement is due to the benchmark corrections; no quantitative effect sizes or sample sizes provided in the excerpt.
Based on these findings, we construct ELT-Bench-Verified, a revised benchmark with refined evaluation logic and corrected ground truth.
Development and release of a revised benchmark (ELT-Bench-Verified) incorporating refined evaluation logic and corrected ground truth as described in the paper.
We develop an Auditor-Corrector methodology that combines scalable LLM-driven root-cause analysis with rigorous human validation (inter-annotator agreement Fleiss' kappa = 0.85) to audit benchmark quality.
Description of a methodology combining LLM root-cause analysis and human validation; human validation reported with inter-annotator agreement Fleiss' kappa = 0.85.
Re-evaluating ELT-Bench with upgraded large language models reveals that the extraction and loading stage is largely solved, while transformation performance improves significantly.
Re-evaluation performed using upgraded LLMs comparing performance across ELT pipeline stages; specific performance metrics or sample sizes not reported in the excerpt.
Constructing Extract-Load-Transform (ELT) pipelines is a labor-intensive data engineering task and a high-impact target for AI automation.
Statement in the paper framing ELT pipeline construction as labor-intensive and high-impact; no empirical data or sample size reported in the provided excerpt.
Applying the Method of Moments Quantile Regression (MMQR) allows the study to capture heterogeneous impacts of robotics across performance levels.
Authors describe use of MMQR in methodology and justify it as appropriate for detecting heterogeneity across quantiles of the dependent variable (value added).
The study uses panel data from Eurostat, the International Federation of Robotics (2024), and World Robotics covering three key sectors in selected EU countries.
Data sources explicitly listed in the paper (Eurostat, IFR 2024, World Robotics); the scope is described as three key sectors in selected EU countries.
Policymakers should support automation through fiscal incentives, invest in reskilling programs, and develop innovation strategies tailored to specific sectors to foster inclusive and sustainable growth.
Policy recommendations derived from empirical findings showing heterogeneous effects of robot density, R&D and human capital across sectors; authors explicitly recommend fiscal incentives, reskilling, and sector-targeted innovation strategies.
The paper’s novelty lies in its differentiated, cross-sectoral approach integrating technological adoption (robotics) with sectoral gross value added using advanced econometric techniques (MMQR).
Authors state the study's contribution is differentiated cross-sectoral analysis and use of MMQR to capture heterogeneous impacts; methodological description provided in paper.
The positive effect of robot density on value added is particularly strong in higher-performing sectors (i.e., at higher quantiles of the value-added distribution).
Results from MMQR showing heterogeneous impacts across performance levels/quantiles; authors state larger positive coefficients of robot density at upper quantiles.
Increased robot density significantly enhances value added.
Empirical analysis using panel data (Eurostat, International Federation of Robotics 2024, World Robotics) estimated with Method of Moments Quantile Regression (MMQR); gross value added used as dependent variable and robot density as a core explanatory variable; authors report statistically significant positive coefficients.
The future of transformative transformer-based AI is fundamentally many, not one.
Concluding synthesis and normative prediction based on the paper's theoretical arguments and literature synthesis; no empirical data or quantified projection provided in the excerpt.
Developing diverse AI teams addresses critics' concerns that current models are constrained by past data and lack the creative insight required for innovation.
Argumentative claim drawing on conceptual critique of current models and the proposed remedy of diverse AI teams; supported by referenced disciplinary literatures but no empirical validation provided in the excerpt.
Having a diverse team broadens the search for solutions, delays premature consensus, and allows for the pursuit of unconventional approaches.
Theoretical/argumentative claim referencing literature in complex systems and organizational behavior as support; no quantitative evidence or sample reported in the excerpt.
Deep intellectual breakthroughs should be expected to come from epistemically diverse groups of AI agents working together rather than singular superintelligent agents.
Predictive/theoretical claim motivated by referenced research and formal results in complex systems, organizational behavior, and philosophy of science; no empirical experiment or sample size given in the excerpt.
We should abandon the individual approach if we're hoping for AI to support groundbreaking innovation and scientific discovery.
Normative prescription based on theoretical argument and synthesis of literature from complex systems, organizational behavior, and philosophy of science; no empirical trial or quantified evaluation reported in the excerpt.
With further development, this approach may exceed traditional methods regarding risk accuracy and help drive innovation in the insurance industry.
Forward-looking claim by the authors extrapolating from current prototype results and potential improvements; no empirical evidence provided that it already exceeds traditional methods.
ARQuest shows great potential to improve user satisfaction and streamline insurance processes.
Interpretation based on experimental findings (fewer questions, user preference) and the proposed framework; forward-looking claim rather than a fully established empirical result.
Adaptive versions were preferred by users for their more fluid and engaging experience.
User preference reported from the experiments (qualitative/user feedback or preference metric); specific measures and sample size not provided in excerpt.
Adaptive versions powered by GPT models required fewer questions.
Experimental result reported in paper comparing question counts between adaptive GPT-powered questionnaires and traditional questionnaires; no numeric counts or sample sizes provided in the excerpt.
Techniques such as social media image analysis, geographic data categorization, and Retrieval Augmented Generation (RAG) are used to extract meaningful user insights and guide targeted follow-up questions.
Described methods/techniques used within the ARQuest system implementation in the paper.
The ARQuest framework introduces a new approach to underwriting by using Large Language Models (LLMs) and alternative data sources to create personalized and adaptive questionnaires.
Methodological contribution described in the paper (framework design); description of components and intended function rather than a quantified outcome.
Achieving near-perfect success rates at this minimally sufficient quality level or comparable success rates at superior quality would require several additional years.
Authors' forecast/commentary on timeline beyond the 2029 projection; conditional expectation based on historical pace of improvements.
If recent trends in AI capability growth persist, LLMs will be able to complete most text-related tasks with success rates of, on average, 80%-95% by 2029 at a minimally sufficient quality level.
Longer-term projection contingent on continuation of recent capability growth trends (model-based forecast stated by the authors).
AI success rates for those tasks increase to about 65% by 2025-Q3.
Short-term projection / trend extrapolation reported in the paper (from the ongoing evaluation data).
In 2024-Q2, AI models successfully complete tasks that take humans approximately 3-4 hours with about a 50% success rate.
Empirical measurement/estimate from the ongoing evaluation (reported temporal snapshot for 2024-Q2); based on tasks mapped to human completion time and observed model success rates from the >17,000 evaluations.
AI performance is high and improving rapidly across a wide range of tasks.
Empirical results from the ongoing evaluation of >3,000 tasks and >17,000 evaluations showing high and increasing success/performance metrics.
Substantial evidence that rising tides are the primary form of AI automation.
Patterns observed in the same large-scale evaluation across tasks and human judgments indicating broad-based, continuous capability improvements across many tasks.
Only interventions that reshape risk allocation can plausibly shift stable system-level behaviour.
Argument based on the paper's game-theoretic reasoning and stylised example (theoretical claim; no empirical testing reported in the abstract).
Artificial intelligence (AI) is widely promoted as a promising technological response to healthcare capacity and productivity pressures.
Author assertion in the paper's introduction/abstract, based on literature/policy discourse (no empirical sample or quantitative analysis reported in the abstract).
Improvements in operational resilience enhance firms' capacity for sustainable development.
Further analysis in the paper showing a positive relationship between OR improvements and indicators of firms' sustainable development capacity.
The enabling effect of AI on operational resilience is more pronounced for capital-intensive enterprises.
Heterogeneity/subsample analysis showing larger AI effects on OR for capital-intensive firms.
The enabling effect of AI on operational resilience is more pronounced for technology-intensive enterprises.
Heterogeneity/subsample tests reported in the paper indicating stronger AI effects on OR for technology-intensive firms.
The enabling effect of AI on operational resilience is more pronounced for enterprises in the growth stage.
Heterogeneity/subsample analysis showing larger AI-induced OR gains among firms classified as in the growth stage.
The enabling effect of AI on operational resilience is more pronounced for enterprises located in the coastal eastern region.
Heterogeneity/subsample analysis reported in the paper showing larger AI effects for firms in the coastal eastern region compared to other regions.
AI promotes operational resilience by optimizing supply chain allocation performance.
Mechanism tests in the paper linking AI adoption to improved supply chain allocation/performance metrics, which are associated with higher OR.
Application of AI significantly enhances corporate operational resilience (OR).
Staggered DID estimation exploiting AIIAPZ policy as quasi-natural experiment on Chinese A-share listed manufacturing firms (2012–2023); main regression results reported as significant.
The empirical results are robust across parallel trend analysis, placebo tests, propensity score matching (PSM), and alternative measures of sustainable performance.
Reported battery of robustness checks listed in the abstract (parallel trend, placebo, PSM, alternative outcome measures).
The R&D deduction policy has stronger effects on larger-scale firms.
Heterogeneity analysis reported in the paper showing larger estimated effects for firms of larger scale.
The R&D deduction policy has stronger effects on non-state-owned firms.
Heterogeneity analysis contrasting policy effects between state-owned and non-state-owned firms reported in the paper.
The R&D deduction policy has stronger effects on firms with high capital intensity.
Heterogeneity analysis in the paper showing larger estimated policy effects for high capital intensity firms.
The R&D deduction policy has stronger effects on firms characterized by rapid technological obsolescence.
Heterogeneity analysis reported in the paper comparing treatment effects across firms with different rates of technological obsolescence.
The policy effect operates by improving total factor productivity (TFP).
Mechanism analysis showing a positive association between the R&D deduction policy and firms' estimated TFP.
The policy effect operates by boosting firms' innovation capabilities.
Mechanism analysis in the paper linking the R&D deduction policy to measures of innovation capability (e.g., innovation output/indicators).
The policy effect operates by alleviating financing constraints for firms.
Mechanism analysis reported in the paper (mediation/heterogeneity analyses linking policy to reduced financing constraints).
The additional deduction policy for R&D expenses (the R&D policy) significantly enhances the sustainable development outcomes of intelligent manufacturing enterprises.
Panel data from listed manufacturing firms in China analyzed using a quasi-natural experiment design; main empirical specification shows a statistically significant treatment effect (abstract reports significance). Robustness checks reported.
HEWU is designed to become the cited standard before better-resourced players define competing frameworks, establishing measurement infrastructure for the cognitive industrial revolution the way GAAP established it for capital markets.
Aspirational/strategic claim made by the authors about intended role and adoption of HEWU (no empirical support provided).