Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The decentralizing effect of digitalization is stronger for firms operating in environments of higher uncertainty.
Moderation analyses in the paper using public data on China's listed companies (2009–2020) showing interaction between digitalization and environmental uncertainty on decentralization outcomes.
The decentralizing effect of digitalization is more pronounced for companies with greater business diversification.
Moderation tests reported in the study using the same dataset of China's listed companies (2009–2020) examining interaction between digitalization and business diversification on subsidiary empowerment.
Firms with higher levels of digitalization tend to decentralize decision‑making authority to their subsidiaries.
Empirical analysis using public data from China's listed companies between 2009 and 2020; paper reports multiple measures of digitalization and tests the relationship between firm digitalization and subsidiary empowerment (decentralization).
The paper documents best practices for iteratively generating tests to capture existing system behavior before model-assisted refactoring.
Methodological contributions in the paper: recommended workflow and practices for iterative test generation to lock down behavior prior to refactoring.
The described workflow constrained refactoring changes and enabled model-assisted refactoring under developer supervision, with proposed code changes validated by passing tests.
Methodological description in the paper: iterative test generation to capture existing behavior, then model-assisted refactoring with developer oversight and test-based validation.
The generated tests achieved up to 78% branch coverage in critical modules.
Measured branch coverage reported in the case study for critical modules after running the generated tests.
Using coding models, we generated nearly 16,000 lines of reliable unit tests in hours rather than weeks.
Single case study reported in the paper: automated unit test generation using coding models; reported aggregate output of generated tests and a qualitative time comparison (hours vs weeks).
The results confirm the expediency of concentrating government support on a limited number of industries with the greatest potential for structural transformation.
Policy recommendation derived from the simulated outcomes of the integrated model showing larger productivity, adoption, and employment gains in the identified priority sectors (model calibrated to 2020–2024 Kazakhstan industry data).
Across all sectors, there is a steady excess of the effect of creating new tasks over the effect of automation (i.e., task-creation effects exceed automation effects), reflecting the specifics of a resource-dependent economy with a shortage of qualified personnel.
Model decomposition of task-creation versus automation effects in the Acemoglu–Restrepo task-oriented component of the integrated model, calibrated with 2020–2024 industry data for Kazakhstan.
Total employment increases by 22.4 p.p. (equivalent to +1.3 million jobs) over the simulation period.
Simulated employment outcomes from the integrated dynamic model calibrated to Kazakhstan industry data (2020–2024).
The share of priority industries in the GDP structure increases by 6.3 p.p. (over the simulation horizon).
Modelled structural transformation outcomes from the integrated simulation using Kazakhstan industry data (2020–2024).
AI adoption in priority sectors by 2035 exceeds the indicators of non-priority industries by 13–32 p.p.
Comparison of simulated AI adoption trajectories between priority and non-priority sectors using the integrated model calibrated on 2020–2024 industry data.
The level of AI adoption in priority sectors reaches 86.8–93.8 p.p. by 2035.
Projection from the paper's Bass-model-based diffusion and integrated dynamic model calibrated to Kazakhstan industry data (2020–2024).
Of the cumulative 35.3 p.p. increase in gross value added for 2025–2035, 16.8 percentage points are attributable to AI.
Decomposition of the simulation results produced by the paper's integrated dynamic model, using industry data for 2020–2024 for calibration.
The cumulative increase in gross value added in the analysed industries will amount to 35.3 p.p. for the period 2025-2035.
Simulation results from the paper's integrated dynamic model (Bass diffusion + expanded production function with endogenous technological progress + task-oriented Acemoglu–Restrepo approach) calibrated to industry data from the Bureau of National Statistics of the Republic of Kazakhstan for 2020–2024.
This work demonstrates how energy considerations can be embedded directly into AI-assisted coding workflows, supporting developers as they engage with energy implications through actionable feedback.
Concluding claim based on the system implementation and evaluation described (benchmarks and controlled study).
EcoAssist reduced per-website energy by 13-16% on average.
Reported result from the benchmark evaluation of 500 websites (effect size reported as 13-16%).
We introduce EcoAssist, an energy-aware assistant integrated into an IDE that analyzes AI-generated frontend code, estimates its energy footprint, and proposes targeted optimizations.
Description of the system introduced by the authors (implementation claim).
AI assistance improves short-term performance on tasks (people do better while using the AI).
Randomized controlled trials (N = 1,222) showing better immediate task outcomes when participants used AI assistance.
Analyses use fixed-effects regression and structural equation modeling (SEM) on panel data from OECD countries.
Methods statement in the paper indicating use of fixed-effects and SEM applied to OECD-country panel data.
This paper provides the first cross-country empirical validation of AI-augmented scientific evaluation systems.
Authors' stated novelty claim that prior work lacked cross-country empirical quantification and that their OECD panel study is the first such validation.
A one standard deviation increase in AIRC is associated with an 18–25% increase in scientific productivity.
Reported point estimate/range from regression/SEM results linking a 1 SD change in the constructed AIRC to productivity outcomes in the OECD panel.
AI-assisted evaluation significantly enhances scientific productivity.
Fixed-effects regression and structural equation modeling (SEM) applied to panel data from OECD countries; reported association between AIRC and research output.
We construct a novel AI Review Capability Index (AIRC).
Paper reports creation of a new composite index (AIRC) to measure national-level AI capability in peer review; constructed and applied to panel data from OECD countries.
Ablation experiments and scalability analysis verify the effectiveness of each core module of HGA-MADDPG.
Ablation study and scalability analysis reported in the paper; experiments removing or altering core modules and reporting comparative performance.
HGA-MADDPG maintains a cost reduction rate of 21.5% in a 120-node ultra-large-scale supply chain.
Scalability experiments reported in the paper on a 120-node simulated supply chain; reported cost reduction rate of 21.5% for HGA-MADDPG.
In the same extreme scenario of triple perturbation, HGA-MADDPG achieves a recovery time of 58 hours, outperforming existing methods.
Simulation experiments under triple perturbation reported in the paper; reported recovery time of 58 hours and stated superior performance relative to baselines.
In an extreme scenario of triple perturbation, HGA-MADDPG achieves a cost deviation rate of 29.6%, which is significantly better than existing methods.
Simulation experiments under an extreme scenario (triple perturbation) reported in the paper; comparison with existing methods and reported cost deviation rate of 29.6%.
In the same baseline scenario, HGA-MADDPG controls the stockout rate at 3.2%.
Simulation experiments reported in the paper (baseline four-level supply chain using real data), reporting a stockout rate of 3.2% for HGA-MADDPG compared to baselines.
In the same baseline scenario, HGA-MADDPG achieves a service level improvement rate of 42.8% compared with eight baseline algorithms.
Simulation experiments reported in the paper (baseline four-level supply chain using SCDL and WSN data), compared to eight baselines; reported 42.8% service level improvement.
In a baseline scenario (four-level supply chain, dynamic environment driven by real data from SCDL and WSN) and compared with eight baseline algorithms, HGA-MADDPG achieves a total cost reduction rate of 26.2%.
Simulation experiments reported in the paper: four-level supply chain baseline scenario driven by real data (SCDL and WSN), compared to eight baseline algorithms; reported aggregate result of 26.2% total cost reduction.
The paper constructs an adversarial disturbance and resilient training architecture that models three types of disturbances (demand mutation, node failure, transportation delay), adversarial agent injection, a dynamic environment replay buffer, and a two-stage training strategy.
Methodological description and implementation details of the training architecture and disturbance models in the paper.
An adaptive fusion weight based on marginal returns is designed to dynamically balance local and global credit.
Methodological description (design and incorporation of adaptive fusion weight in algorithm).
The algorithm quantifies the contribution of individual actions to sub-chain objectives and system-level indicators through local and global credit networks.
Methodological description and algorithm design (local and global credit networks described in the paper).
HGA-MADDPG introduces a hierarchical graph attention mechanism to dynamically represent the state of the supply chain network topology.
Methodological description and algorithm design presented in the paper (development and implementation of the hierarchical graph attention mechanism).
All code, infrastructure, and benchmark data are released to facilitate future research in realistic computer-use agents.
Statement of release in paper (availability claim).
Applying the same auditing principle at test time — a separate VLM reviews completed trajectories and provides feedback — improves Gemini-3-Flash on CUA-World-Long from 11.5% to 14.0%.
Experimental result reported in paper: evaluation of Gemini-3-Flash with/without test-time VLM auditing on CUA-World-Long, reported scores 11.5% -> 14.0%.
Distilling successful trajectories from the training split into a 2B vision-language model outperforms models 2× its size.
Modeling experiments reported in paper: distilled 2B VLM evaluated against larger models (2× size). Exact evaluation metrics and baseline model sizes not specified in excerpt.
CUA-World-Long is a challenging long-horizon benchmark with tasks often requiring over 500 steps, far exceeding existing benchmarks.
Benchmark description in paper reporting typical task lengths ("often requiring over 500 steps") and comparison to existing benchmarks.
The result is CUA-World, a collection of over 10K long-horizon tasks spanning domains from medical science and astronomy to engineering and enterprise systems, each configured with realistic data along with train and test splits.
Dataset release / creation claim specifying >10,000 tasks and train/test splits.
Using a taxonomy of economically valuable occupations grounded in U.S. GDP data, we apply this pipeline to 200 software applications with broad occupational coverage.
Dataset creation procedure and reported coverage claim (200 software applications), taxonomy derived from U.S. GDP data as stated.
Environment creation is framed as a multi-agent task: a coding agent writes setup scripts, downloads real-world data, and configures the software while producing evidence of correct setup; an independent audit agent verifies evidence against a quality checklist.
Method description of multi-agent pipeline (coding agent + audit agent) in the paper.
We introduce Gym-Anything, a framework for converting any software into an interactive computer-use environment.
Methodological contribution described in paper (framework implementation claimed).
SWE-bench alignment: Bench is aligned with SWE-bench-Verified and SWE-bench-Pro.
Paper statement that the constructed benchmark is aligned with SWE-bench-Verified and SWE-bench-Pro (methodological/design alignment described).
Bench contains 495 issues and 1,787 validated design constraints across six repositories.
Reported dataset statistics in paper/abstract: explicit counts of issues (495), validated constraints (1,787), and number of repositories (6).
We construct DESIGN-AWARE benchmark (Bench) by mining and validating design constraints from real-world pull requests, linking them to issue instances, and automatically checking patch compliance using an LLM-based verifier.
Method description in paper: dataset created by mining real-world pull requests, validating constraints, linking constraints to issues, and using an LLM-based verifier to check compliance.
Flowr is domain-independent, offering a generalizable blueprint for agentic AI-driven supply chain automation across large-scale enterprise settings.
Claim of generalizability made by the authors in the paper; presented as an assertion rather than demonstrated through multi-industry empirical tests in the excerpt.
The framework was validated in collaboration with a large-scale supermarket chain.
Claim of field validation stated in the paper; indicates at least one real-world collaboration but provides no further details (e.g., number of stores, duration, metrics) in the excerpt.
Evaluation indicates Flowr enables proactive exception handling at a scale unachievable through manual processes.
Empirical/operational claim based on the paper's evaluation and deployment context; the excerpt asserts this capability but does not provide quantitative performance metrics or comparison details.
Evaluation shows Flowr improves demand–supply alignment.
Empirical claim in the paper's evaluation; reported improvement in demand-supply alignment from deployment or testing with a large supermarket chain, but no numerical metrics provided in the excerpt.