Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The positive effect of industrial robots on firm-level TFP is statistically significant among firms that face low government subsidies.
Heterogeneity/subsample analysis in the paper using firm-level data (2006–2019) that splits or interacts robot effects with measures of government subsidies.
The application of industrial robots has a positive effect on firm-level total factor productivity (TFP).
Empirical analysis using matched data on industrial robots and Chinese listed companies for 2006–2019; the paper reports econometric tests estimating the relationship between robot application/adoption and firm-level TFP.
Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.
Synthesis of empirical results (agreement, low accuracy vs. experts, and user-satisfaction models) leading to design recommendations discussed in the paper.
Proactive interactions (by the system) positively affect user satisfaction.
Same user-satisfaction model reported in the paper showing positive association between proactive system behavior and satisfaction across participant sessions.
Developers tend to agree with LLM assessments of NFRs.
Empirical analysis of the 49 programmers' interactions and judgments compared to the LLM's assessments across 148 HIPAA-derived NFRs (agreement rates computed in the paper).
Higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing (Capability Frontier) and the best single model (shown via controlled probabilistic simulations).
Controlled probabilistic simulations reported in the paper; demonstration of relationship between query topic entropy and performance gap between oracle routing and best single model. (Simulation parameters not specified in the excerpt.)
Using the Capability Frontier, state-of-the-art (SOTA) accuracy is matched at an 85% cost reduction.
Empirical cost-performance comparison across the 21 LLMs and 16 benchmarks showing the Capability Frontier achieves SOTA accuracy at substantially lower cost.
Additionally correcting for single runs (i.e., sampling multiple generations and selective retention) yields an 82% improvement.
Empirical comparison using the Capability Frontier with both corrections (across 21 LLMs and 16 benchmarks) versus single-run, single-model benchmark evaluations.
Correcting for single-model evaluation yields a 54% error rate reduction.
Empirical evaluation comparing the Capability Frontier to each benchmark's top-performing model at matched cost across 21 LLMs and 16 benchmarks (coding, reasoning, medicine, factuality, instruction following, agentic).
AI-driven agglomeration can support green agricultural transformation.
Overall interpretation and policy implication derived from empirical findings (U-shaped relationship, mediation by technological progress, heterogeneity and threshold results) based on the provincial panel analysis.
The dynamic (U-shaped) process is more likely to emerge when public innovation investment and rural household income exceed critical thresholds.
Threshold analysis reported in the paper indicating the nonlinear relationship becomes apparent when public innovation investment and rural household income pass estimated threshold values, based on the 30-province panel.
Agricultural socialized services and rural industrial integration buffer the initial negative association between AIIA and APCRS.
Moderation/interaction analysis in the paper showing those regional/institutional factors mitigate the early negative effect of AIIA on APCRS using the 30-province panel.
Technological progress partially mediates the relationship between AIIA and APCRS.
Mediation analysis reported in the paper using the same 30-province panel (2016–2024) indicating a partial mediation effect of technological progress.
Together, these contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.
Synthesis of the paper's contributions (case study, new benchmarks, empirical analyses, randomized experiment).
The observed ~2x speedup is likely underestimated because one-fifth of human-only reproductions reached the time limit before completing.
Observed censoring/timeouts in the randomized experiment (authors report ~20% of human-only runs hit the time limit).
The randomized experiment found a statistically significant speedup by about a factor of two for human-agent collaboration compared to human-only on computational reproducibility tasks.
Small-scale randomized experiment reported in the paper (statistical test reported as significant; exact n not stated in the abstract).
Despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring reliability and model versus scaffold performance.
Empirical analyses on CORE-Bench v1.1 separating model contribution from scaffold/infrastructure contribution and reporting reliability-related metrics.
Despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency.
Empirical measurements on CORE-Bench v1.1 reported in the paper (benchmark timing/efficiency analyses).
We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD.
Development and release of two benchmark artifacts described in the paper (methodological contribution).
Measuring agents along the six non-accuracy dimensions yields meaningful insights into agent performance even after accuracy saturates.
Empirical case study using CORE-Bench Hard and the authors' derived benchmarks and analyses.
For Chinese firms, productivity gains from GenAI are most likely when adoption is supported by cloud infrastructure, data readiness, skilled labor, workflow redesign, and strong digital ecosystems.
Synthesis of China-focused digital transformation studies and literature incorporated into the review; no new China-specific empirical analysis in this paper.
Existing studies show that GenAI can improve software-development tasks.
Synthesis of empirical studies and task-level experiments (e.g., developer assistance tools) reviewed in the paper.
Existing studies show that GenAI can improve consulting tasks.
Cited task-level studies and applied examples in advisory/consulting work synthesized in the review.
Existing studies show that GenAI can improve customer support tasks.
Review of task-level experiments and applied studies in customer support settings reported across the literature synthesized in the paper.
Existing studies show that GenAI can improve writing tasks.
Synthesis of task-level productivity experiments and prior empirical studies on GenAI-assisted writing (literature reviewed in the paper).
AIGP simultaneously provides interpretable and transparent pricing rationales.
Author claim in paper that the LLM-based approach produces interpretable pricing rationales; no details or quantified user/analyst evaluation provided in the excerpt.
In large-scale online A/B tests on Tao Factory, AIGP achieved +8.20% in milestone achievement rate over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
In large-scale online A/B tests on Tao Factory, AIGP achieved +7.59% in Return on Investment (ROI) over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
In large-scale online A/B tests on Tao Factory, AIGP achieved +13.21% in Gross Merchandise Value (GMV) over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
The Long-Term Value Estimator (LTVE), trained via offline reinforcement learning on historical data, serves as a reward model to score candidate pricing actions and select preference pairs for Direct Preference Optimization (DPO), thereby aligning the pricing policy with long-term business objectives.
Methodological description in paper: LTVE trained by offline RL on historical data and used as reward model for scoring and DPO preference selection; described as central component of framework. No sample size or quantitative validation details in the excerpt beyond its described role.
For efficient deployment while maintaining high-quality outputs, AIGP employs supervised fine-tuning for knowledge distillation.
Methodological claim in paper describing use of supervised fine-tuning (SFT) / knowledge distillation as part of model training and deployment strategy; no quantitative performance numbers tied specifically to SFT in the provided excerpt.
AIGP is a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual context to make interpretable, knowledge-aware pricing decisions.
Methodological description in the paper (proposed system architecture/design); no quantitative evaluation detail provided in the excerpt beyond assertions.
AI-driven outcomes depend less on the technology itself and more on complementary conditions—human capital formation, digital and data infrastructure, institutional coordination, and governance capacity—that enable effective diffusion.
Thematic synthesis of reviewed literature (2015–2025) highlighting repeated findings that complementarities (human capital, infrastructure, institutions) mediate AI diffusion and impacts.
The research aims to inform the design of responsible sociotechnical systems that preserve meaningful human involvement while retaining the efficiency benefits of AI-enabled decision support.
Stated normative aim/purpose in the abstract (intended policy/design impact). No empirical evidence or evaluation reported in the provided text.
Uncertainty-aware emulation transforms mechanistic crop simulation from a computational bottleneck into an on-demand discovery engine capable of interrogating the full genotype, environment and management space at a scale no process-based model can match.
Synthesis claim based on method development, large training dataset (2 million simulations), speed improvements, and demonstrated large-scale experiments (100k trait configurations).
Radiation use efficiency and temperature-driven root dynamics are dominant drivers of yield resilience.
Analysis of trait importance or sensitivity within the emulator experiments indicating RUE and temperature-driven root dynamics as dominant factors influencing yield resilience.
We identify 181 maize trait combinations that consistently maintain high yield across all tested conditions.
Results of the large-scale emulator exploration across trait configurations, soils, and climate scenarios reporting identification of 181 trait combinations meeting the 'consistently high yield' criterion.
Applying the framework across 100,000 trait configurations, six soil environments in Iowa and Illinois, and climate projections through the year 2100 under two emissions scenarios enables large-scale exploration.
Application/experiment described using 100,000 trait configurations, six soil environments (Iowa and Illinois), and climate projections to 2100 under two emissions scenarios.
The framework provides calibrated predictive uncertainty without costly Bayesian inference.
Modeling approach described as probabilistic, yielding calibrated predictive uncertainty; claim that this is achieved without expensive Bayesian methods.
The framework is augmented with a convolutional synthetic weather generator that produces physically consistent climate sequences.
Methodological statement describing augmentation of the emulator with a convolutional synthetic weather generator intended to generate physically consistent climate sequences.
The emulator was trained on two million simulations spanning diverse genetic, soil, and management conditions.
Method description stating training dataset size of two million APSIM simulations covering varied genetics, soils, and management.
The emulator reduces simulation time by several orders of magnitude compared to the mechanistic APSIM model.
Reported comparison of simulation runtimes between the probabilistic neural emulator and APSIM (statement of 'several orders of magnitude' speedup).
We develop a probabilistic neural emulator of APSIM that reproduces key maize growth processes across 13 outputs with high fidelity (with R^2 of 0.93).
Model evaluation comparing emulator outputs to APSIM across 13 output variables; reported R^2 = 0.93. (Training and test data drawn from simulations used to train the emulator.)
Production scale positively moderates both the human capital mechanism and the product R&D mechanism through which AI promotes value chain upgrading.
Moderation analysis in the panel econometric framework using 30-province data (2010–2022) showing that larger production scale strengthens the mediating effects of human capital and product R&D.
AI facilitates value chain upgrading in the equipment manufacturing industry through two channels: enhancing human capital levels and driving product R&D.
Mechanism tests (mediation/ channel analysis) conducted on the 30-province panel (2010–2022) showing empirical support for human capital and product R&D as mediators of AI's effect on upgrading.
AI significantly enhances value chain upgrading in capital-intensive and technology-intensive equipment manufacturing industries.
Industry-type heterogeneity analysis within the 30-province panel (2010–2022) comparing capital-intensive and technology-intensive subsectors; reported statistically significant positive coefficients for these subsectors.
The positive effect of AI on value chain upgrading remains robust after a series of stability tests and when addressing endogeneity concerns.
Stability/robustness tests and endogeneity discussions reported in the paper applied to the same 30-province panel (2010–2022); unspecified robustness procedures and endogeneity treatments mentioned.
AI promotes value chain upgrading in the equipment manufacturing industry.
Panel econometric analysis using data from 30 Chinese provinces over 2010–2022; models report a statistically significant positive coefficient on AI measures; robustness checks reported.
In June 2026, the median researcher generated more than 50 times as many monthly output tokens across Codex and ChatGPT as they did in November 2025.
Internal monthly output-token counts by role (researcher) from November 2025 and June 2026 in Codex and ChatGPT usage logs.
In June 2026, the median OpenAI employee in a legal role generated 13 times more monthly output tokens across Codex and ChatGPT than they did in November 2025.
Internal monthly output-token counts by role (legal) from November 2025 and June 2026 in Codex and ChatGPT usage logs.