Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Request complexity has increased: since the start of the year, the share of individual Codex users who submit at least one request for a task estimated to require more than eight hours for an experienced human to complete has increased nearly tenfold.
Estimated task-duration analysis applied to user requests in Codex logs, comparing share of users submitting requests classified as >8-hour tasks over time.
26.6% of users use skills, which allow users to share instructions for complex workflows.
Feature-usage statistics from Codex showing proportion of users who invoked or used 'skills' functionality.
More than 10% of users manage three or more concurrent Codex agents at some point each week.
Usage telemetry measuring number of concurrent Codex agents managed by users on a weekly basis.
Within OpenAI, Codex usage is nearly universal and has largely replaced business usage of ChatGPT.
Internal OpenAI employee usage metrics showing high prevalence of Codex and relative decline in business ChatGPT usage among employees.
Agentic AI usage is growing rapidly: the number of active users has grown more than fivefold in the first half of 2026.
Analysis of OpenAI Codex usage logs using an automated, privacy-protecting pipeline comparing active user counts over the first half of 2026.
The framework is operationalised as a Supply Certainty Index (SCI) over five measurable properties, a five-level Determinism Maturity Model (DMM) as an adoption ladder, and an open-question programme (OQ1–OQ5) with explicit null results that would force retraction.
Descriptive / methodological content in the paper: authors define SCI, DMM, and OQ1–OQ5 and describe their components and falsifiable null results.
There is a convergence condition for environment-side skill evolution (a formal result giving conditions under which environment-side skills converge).
Formal theoretical result in the paper (mathematical condition/criterion for convergence of skills evolving on the environment side).
Environment determinism is a complementary binding axis cutting across known frictions (data wall, abstraction barrier, embodied bottleneck, multi-agent trust) for the broad class of agentic AI tasks whose outcomes are verifiable economically, physically, or through multi-party settlement.
Conceptual argument supported by the paper's formal results and framing (authors link the determinism axis to the previously identified frictions and to verifiable outcomes).
The study provides policymakers with a solid empirical foundation for assessing how the diffusion of AI supports inclusive growth and sustainability goals.
Authors' interpretation of their comparative, multi-method cross-sectional findings (factor analysis, GLM, cluster analysis) across EU countries linking AI adoption to economic performance, S&T workforce share, employment, and SDGI.
AI adoption has a weaker but present positive association with sustainability indicators as measured by the Sustainable Development Goals Index (SDGI).
Cross-sectional analysis (factor analysis, general linear models) connecting enterprise-level AI adoption measures to country-level SDGI scores across EU countries; the relationship is described as weaker than for GDP per capita or S&T workforce share; no numeric effect sizes reported in the summary.
AI adoption shows weaker but still present positive relationships with overall employment (total employment) across EU countries.
Cross-sectional general linear model estimations and factor analysis relating enterprise-level AI adoption indicators to total employment across EU countries; the summary states the association is weaker yet present; exact sample size and statistical magnitudes not reported.
AI adoption is associated with a larger share/proportion of highly educated science and technology workers in countries.
Cross-sectional comparative analysis of EU countries using factor analysis and general linear models linking enterprise-level AI adoption measures to the proportion of highly educated science & technology professionals; sample consists of EU countries (exact count not reported).
AI adoption is consistently linked with higher economic performance (GDP per capita) across EU countries.
Cross-sectional analysis across EU countries using exploratory factor analysis and general linear model estimations relating enterprise-level AI adoption indicators to GDP per capita (SD variables: AI adoption, GDP per capita). Sample described as EU countries; exact N not reported in the summary.
AI technologies are increasingly transforming agricultural production, renewable energy generation, waste management efficiency, and environmental sustainability across developing economies.
Background/introductory claim in the study likely supported by literature review and contextual framing rather than primary data from this survey.
The results contribute to emerging literature on AI, renewable energy systems, and sustainable technological development in developing economies and offer practical policy recommendations for Nigeria's green transition agenda.
Authors' stated contribution and recommendations based on study findings from survey (N = 522) and associated analyses.
AI investment significantly improves sustainability outcomes (β = 0.55, p < 0.01).
Multiple regression analysis on survey data (N = 522) reporting coefficient β = 0.55 with p < 0.01 for AI investment predicting sustainability outcomes.
AI investment significantly improves operational efficiency (β = 0.62, p < 0.01).
Multiple regression analysis on survey data (N = 522) reporting coefficient β = 0.62 with p < 0.01 for AI investment predicting operational efficiency.
Positive social impacts (Mean = 3.78).
Survey descriptive statistics (N = 522) reporting mean social impact score = 3.78.
Enhanced energy recovery and environmental sustainability (Mean = 3.95).
Survey descriptive statistics (N = 522) reporting mean score for energy recovery/environmental sustainability = 3.95.
Significant improvements in operational efficiency (Mean = 4.02).
Survey descriptive statistics (N = 522) reporting mean operational efficiency score = 4.02.
Moderate-to-high AI adoption (Mean = 3.84).
Quantitative survey of 522 respondents across Nigeria's six geopolitical zones; descriptive statistics reporting mean adoption score = 3.84.
The main findings (that digital technology adoption improves agricultural production efficiency) are robust to instrumental variable estimation, propensity score matching, and alternative specifications.
Robustness checks reported in the paper: instrumental variable (IV) estimation, propensity score matching (PSM), and alternative model specifications applied to CFPS panel data.
Production efficiency gains significantly enhance farmers' livelihood resilience by strengthening income stability, asset accumulation, and risk-coping capacity.
Further analysis of transmission effects in the paper linking estimated production efficiency gains to livelihood resilience indicators (income stability, asset accumulation, risk-coping capacity) using CFPS panel data.
The effects of digital technology adoption on production efficiency are more pronounced among young and middle-aged farmers, grain producers, households in plain areas, and villages with better digital infrastructure.
Heterogeneity/subgroup analysis using CFPS panel data (interaction or subgroup regressions comparing age groups, crop types, geographic/plain vs other areas, and village digital infrastructure levels).
The productivity effect of digital technology adoption operates primarily through technical efficiency improvement, factor allocation optimization, and service cost reduction.
Mechanism analysis reported in the paper (channel/mediation analysis using CFPS panel data and econometric decomposition of effects).
Digital technology adoption significantly improves agricultural production efficiency.
Empirical analysis using China Family Panel Studies (CFPS) panel data 2014–2022; main estimation and robustness checks reported (instrumental variable estimation, propensity score matching, alternative specifications).
There are two distinct integration pathways through which AI creates supply chain value: a downstream (customer integration → responsiveness) pathway and an upstream (supplier integration → efficiency) pathway.
Synthesis of significant direct effects and moderating effects found via hierarchical regression on survey data from 426 AI-adopting Chinese manufacturing firms.
Supplier integration strengthens (positively moderates) the effect of AI capability on supply chain efficiency.
Hierarchical regression analysis testing moderating effects using survey data from 426 AI-adopting Chinese manufacturing firms.
Customer integration strengthens (positively moderates) the effect of AI capability on supply chain responsiveness.
Hierarchical regression analysis testing moderating effects using survey data from 426 AI-adopting Chinese manufacturing firms.
AI capability significantly enhances supply chain efficiency.
Hierarchical regression analysis on survey data from 426 AI-adopting Chinese manufacturing firms.
AI capability significantly enhances supply chain responsiveness.
Hierarchical regression analysis on survey data from 426 AI-adopting Chinese manufacturing firms.
Multi-model execution-trace-derived skills outperform skills derived from any single-model trace source.
Reported experimental comparisons in the paper between skills evolved from diverse multi-model execution traces and skills from single-model trace sources (statements that multi-model source outperforms all single-model sources).
Skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources.
Experimental comparison in the paper where skills were evolved from multi-model execution traces and evaluated for cross-model test accuracy on the benchmark; reported cross-model test accuracy is 73.1% and stated to outperform single-model trace sources.
Procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points.
Experimental results reported on the AFTER benchmark (experimental evaluation across tasks/roles/models). The paper reports aggregate performance change after one refinement round as 3.7-6.7 points; exact sample breakdown implied to be across benchmark tasks and evaluation settings.
Adding a commit-count control roughly halves Copilot's within-repo co-authorship gap (from +36.2 pp to +24.4 pp).
Regression/adjusted analysis on Copilot PRs including within-repo and commit-count controls; reported before/after point estimates.
Stratifying 33,596 PRs by agent identity reverses the pooled conclusion: Copilot and Devin show large positive within-agent gaps (+41.2 and +33.5 percentage points, both p<0.001), while Cursor, Claude Code, and Codex show small effects whose cross-sectional 95% CIs span zero.
Stratified analysis by agent identity on the AIDev dataset (33,596 PRs); reported within-agent differences and p-values/CI statements.
Researchers must build data infrastructure, adopt participatory methods, and write with policymakers in mind to help close the research–policy gap.
Authors' prescriptive recommendations derived from their analysis of researcher responsibilities in bridging the gap.
Closing the research–policy gap requires policymakers to widen their evidence base, engage workers as epistemic partners, and shift from prediction to preparedness.
Authors' prescriptive recommendations based on their analysis of the coordination gap.
Five families of research are responding to these limits: dynamic and benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption and usage data.
Survey of related literature and emerging methods described in the paper.
Eloundou et al. (2023) define exposure as the share of occupational tasks a large language model can assist with (the 'GPTs are GPTs' or GPTs scores).
Direct definitional citation of Eloundou et al. (2023) provided in the paper.
A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate.
Statement in paper referencing the widespread use of exposure scores produced in Eloundou et al. (2023) as an input to policy and research discussions.
The employment-enhancing mechanisms encompass productivity, real income growth, complementary jobs, new jobs and sectors, market expansion and commodification.
Mechanisms listed by the author as explanatory pathways, drawn from the paper's comprehensive theoretical and empirical literature review (no empirical quantification provided in the excerpt).
In countries undergoing intensive automation, there has been a rapid increase in the number and proportion of workers, rather than a decline.
Asserted as an empirical finding derived from the paper's review of studies of countries with intensive automation; the excerpt does not list specific countries, datasets, or sample sizes.
The employment-enhancing effects of new technologies are demonstrated.
Stated as a conclusion based on a 'comprehensive review of theoretical and empirical studies' (no specific studies, sample sizes, or quantitative meta-analytic statistics reported in the excerpt).
GAIE explicitly bridges AI-assisted development maturity with regulatory governance through proportionate human oversight.
Conceptual claim about the framework's contribution as presented in the paper (theoretical contribution).
Graduated oversight preserves compliance evidence coverage for regulated functions.
Regulatory coverage analysis and framework specifications in the paper (qualitative/analytic demonstration rather than empirical validation).
Evaluation through regulatory coverage analysis, comparative framework analysis, and analytical productivity modeling suggests that graduated oversight preserves 84--97% of agentic coding velocity (central estimate: 91%) while maintaining compliance evidence coverage for regulated functions.
Analytical productivity modeling combined with regulatory/comparative analyses described in the paper (modeling results; no reported empirical trial/sample size).
We map GAIE against the Bank of Thailand's 2025 AI risk-management policy and demonstrate cross-jurisdiction applicability to MAS (Singapore), NIST AI RMF, ISO/IEC 42001, and the EU AI Act.
Regulatory coverage analysis and mapping exercise described in the paper (comparative policy analysis).
Each oversight tier defines required evidence artifacts for compliance auditability.
Framework design specifying evidence/artifact requirements per tier (methodological specification in the paper).
GAIE introduces the Oversight Classification Model (OCM), a deterministic decision function that classifies code generation tasks by regulatory impact, customer proximity, reversibility, and data sensitivity to route them through one of three oversight tiers: human-in-the-loop (strategic functions), human-over-the-loop (customer-impacting), or automated-with-monitoring (internal).
Specification of the OCM within the framework (model/method description in the paper).