Evidence (2488 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
A one-time reduction in the cost of building generative-AI capability revives the build route primarily before the leading design has stabilized; after the dominant design is established, the effect is substantially weaker.
Agent-based model experiment varying the timing of an open-weight-style reduction in build costs relative to dominant-design formation; stated as Proposition P3.
The paper distinguishes substantive peer competition from symbolic peer competition by measuring the former with peers’ real digital-transformation investments and the latter with peers’ disclosures or announcements.
Operationalization of the two peer-competition measures in the empirical analysis.
Funding mandates, international collaboration, journal prestige, and disciplinary norms are associated with systematic differences in Creative Commons license selection.
The study models categorical CC-license choice using multinomial logistic regression and examines funding-policy strength, international collaboration, subject category, and journal impact-factor percentile among 122,085 open-access articles.
The two banks studied use different primary performance metrics: one relies on Unit Profitability, while the other has begun adopting Customer Profitability.
Comparative qualitative case study of two national commercial banks, using document analysis and stakeholder insights.
AI adoption and performance in hospitality-like settings should be expected to vary across managers because blended strategic-operational roles, managerial adaptability, leadership style, and governance context can produce heterogeneous implementation choices and returns.
Conceptual implication derived from the hospitality GM framework; no direct AI adoption experiment or quantified treatment-effect estimate is reported.
In Vietnam, AI adoption is uneven and is concentrated primarily among large enterprises, financial institutions, and technology firms with relatively advanced digital infrastructure and financial resources.
Qualitative contextual analysis based on secondary sources and comparison of global best practices with Vietnam's adoption conditions; no representative enterprise survey is reported.
Organizational capacity, technological fit, perceived advantage, training, and organizational design may matter for adoption, but their influence depends partly on the degree of political support and legitimacy given to implementation efforts.
Qualitative interview findings and the paper's theoretical argument concerning political authorization in public-sector technology implementation.
Among more than 120,000 workers tracked quarterly, 82% of AI users sustain usage quarter over quarter, but only about 2% reach a level of use consistently embedded in their workflows.
The paper cites ActivTrak quarterly tracking of more than 120,000 workers.
In a 28-day longitudinal study, daily conversations with AI shifted participants’ preferences toward AI and away from humans when the conversations became personal.
Longitudinal study conducted with OpenAI in which participants engaged in daily five-minute AI conversations over approximately one month across personal, non-personal, and open-ended topics; future preferences for human versus AI support were measured.
Technical feasibility is necessary for agentic adoption but is not sufficient to explain where agentic adoption occurs.
This conclusion is based on the observed positive alignment between AAI and technical capability measures at lower levels of adoption, combined with the unexplained adoption shortfall among high-wage and highly educated occupations.
Technical availability explains most of the variation in agentic adoption at the low end, but does not explain the adoption shortfall among the most educated occupations.
The authors compare AAI with technical-exposure measures and observe that highly educated occupations rank high on technical exposure while adopting agents relatively little.
The AAI peaks among occupations requiring a bachelor’s degree and declines among occupations at both lower and higher education extremes.
The authors relate occupation-level AAI scores to typical entry-level education requirements from the 2025 BLS Occupational Employment and Wage Statistics data.
Agentic adoption follows an inverted-U relationship with occupational wages, with adoption peaking below the top of the wage distribution.
The authors examine AAI variation by occupational median wage using 2025 Bureau of Labor Statistics wage data for occupations with available wage and education information.
The occupations where agent adoption concentrates differ sharply from the occupations identified by earlier automation research as most exposed.
The authors compare the Agentic Adoption Index (AAI) with the pre-AI probability-of-computerisation measure and other occupational AI-exposure measures across 748 O*NET occupations.
For Saudi SMEs, cloud adoption is associated with perceived benefits, organizational factors, security concerns, expertise, and provider-related conditions.
Saudi-focused adoption studies cited by the review, particularly Alamri and Alzahrani (2024), along with Alqahtani et al. (2023).
Cloud adoption is influenced by combinations of technological, organizational, and environmental conditions rather than by a single universal factor.
Synthesis of cloud-adoption research, including configurational analysis by Zhang et al. (2021), and studies identifying security, cost, compatibility, readiness, leadership, skills, and external conditions.
Saudi Arabia has strong institutional and infrastructural conditions for adoption of big data, AI, and ML, but talent shortages, fragmented legacy data, model risk, privacy requirements, and uneven small-firm readiness remain major constraints.
Review synthesis of Saudi national policy, digital infrastructure, localization objectives, industrial ecosystems, Saudi organizational research, and SME research.
The benefits and outcomes of AI adoption in Asian firms are moderated by institutional and regulatory regimes, digital infrastructure, data quality, human capital, financing constraints, firm size, and sector.
Conceptual discussion of Asian context and moderating factors; the paper does not estimate moderator effects using cross-country or firm-level data.
EU countries exhibit substantial heterogeneity in digital integration, organizational innovation, and AI-adoption intensity.
Comparative Eurostat indicators analyzed using K-means cluster analysis.
AI-related job shares are highly unequal across US counties: Slope County, North Dakota, had an AI job share of 10%, Santa Clara County, California, had a share of 8.2%, and many rural counties had virtually none.
Andreadis et al. analyzed US county-level job-posting data from 2014 to 2023.
Willingness to send agent-mediated communication and willingness to receive or engage with agent-mediated communication are distinct constructs, despite being highly correlated.
Two-dimensional graded response model with latent regression applied to survey responses; model comparison favored separate send and receive dimensions over a single construct.
The effect of liability on AI use for disadvantaged patients is non-monotone: use initially declines as liability increases, but can rise at higher liability levels.
Equilibrium analysis of the linked firm-design and physician-use model. At low liability, direct deterrence dominates; at higher liability, the firm's endogenous accuracy investment and possible switch to equal accuracy reduce physician exposure.
Early adoption of military AI may provide operational advantages while simultaneously increasing exposure to exploitation of immature systems and networks.
Conceptual first-mover-versus-fragility analysis; no empirical estimate or formal model is presented.
Military AI competition creates incentives for rapid adoption and integration of autonomous weapons, ISR systems, and decision-support tools, while also generating negative externalities such as security races and underinvestment in verification and resilience.
Interpretive analysis of adoption pressures and investment incentives, combined with the paper's economic implications regarding externalities and defensive investment.
Whether firms adopt generative AI for a task depends not only on technical exposure but also on verification costs, liability, trust, and governance conditions.
Theoretical adoption condition that explicitly incorporates verification and governance frictions into the decision to deploy AI.
The paper’s measure of organizational generative-AI adoption captures formalized hiring requirements and stated work practices rather than all informal employee experimentation.
Measurement constructed from Lightcast labels identifying generative-AI tools or skills mentioned in job-posting text; the paper notes that informal use may not be observed.
The available public evidence shows Indian benchmark participation concentrated in foundational coding, while it remains absent from the reviewed agentic coding benchmarks.
Comparison of Indian reporting on HumanEval, MBPP, and LiveCodeBench with the absence of Indian results on SWE-bench Pro, DeepSWE, Terminal-Bench 2.1, and Vibe Code Bench v1.1.
The dataset should be interpreted as a lower bound on the full population of agent skills because it covers public repositories only and GitHub code search excludes some files and repositories.
The authors document restrictions to public repositories, default branches, files under 384 KB, recently active repositories, and selected forks.
In the monopoly model, improvements in AI's complex-task capability increase AI adoption, whereas improvements in easy-task capability may reduce adoption because the provider optimally rations access toward users with high willingness to pay.
Lemma 1 establishes that the optimal access cutoff is increasing in complex-task capability bd,t and decreasing in easy-task capability be,t, under a uniform distribution and an interior cutoff.
Estimated LLM usage in 2025 was higher for papers from predominantly non-native-English-speaking countries than for papers from predominantly native-English-speaking countries: 72% versus 37%.
Grouping papers by author-affiliation countries and comparing the estimated LLM usage for the two country groups.
Among the countries with the largest numbers of papers in the dataset, estimated full-paper LLM usage in 2025 was highest in South Korea at 85%, followed by China at 82% and Taiwan at 80%, and lowest in the UK at 28%.
Country-level estimates for the 20 countries with the largest paper counts, using a fixed set of marker words rather than country-specific optimization.
Over the complexity range between the platform's and consumer's thresholds, agentic search would generate higher platform revenue if used, but the consumer rationally continues manual search.
Analytical threshold-wedge result based on the platform maximizing conversion revenue and the consumer minimizing mismatch plus search expenditure.
If long-run adoption reaches only 80% to 100% of codifiable tasks, the model produces a peak in approximately half of simulation draws; below that range, a peak does not occur.
Monte Carlo analysis drawing the long-run adoption ceiling from a uniform distribution on 0.8 to 1.0, combined with the peak-coverage threshold.
The study challenges the assumption that executive psychological resilience is uniformly beneficial for enterprise digital transformation.
The reported inverted-U finding shows that resilience has positive effects at moderate levels but negative effects beyond an optimum.
The resilience level that maximizes enterprise digital transformation is intermediate rather than maximal.
The reported nonlinear relationship identifies an optimal resilience level at the vertex of the inverted-U specification.
Executive psychological resilience has an inverted U-shaped relationship with enterprise digital transformation in Chinese A-share listed firms from 2011 to 2024: moderate resilience promotes digital transformation, while excessively high resilience reduces transformation effectiveness.
Panel-data analysis of Chinese A-share listed firms using a nonlinear quadratic specification relating executive psychological resilience to enterprise digital transformation.
Singapore attracts over 75% of Southeast Asia's AI venture-capital funding, while Indonesia and Vietnam have population-level AI adoption rates of approximately 42%.
Regional comparison attributed to the cited source [12]; the chapter uses these figures to illustrate a decoupling between policy readiness, investment concentration, and population-level adoption.
Large organizations are more likely than smaller organizations to integrate AI into their work processes; only 20% of organizations with fewer than 1,000 workers were reported as integrating AI.
Organizational-size comparison from the cited ASEAN business adoption evidence.
Although 85% of surveyed Southeast Asian businesses reported using AI in their organizations, more than 80% remained in the initial phases of their AI journey because organization-wide integration into production processes or workflows was still limited.
ASEAN Secretariat survey/report cited by the chapter; the chapter distinguishes organizational AI use from full workflow or production integration.
AI adoption is uneven among Chinese listed firms: the median value of the AI application indicator is zero.
The study measures AI application using the log-transformed frequency of AI-related terms in annual reports and reports descriptive statistics for 38,190 firm-year observations.
Practitioner reports supply longitudinal and cross-institutional adoption statistics but generally rely on self-reported institutional surveys and descriptive analyses.
Methodological characterization of the 8 practitioner and policy reports.
Peer-reviewed research is substantially richer on micro-credential purposes, enabling conditions, and barriers than on measured adoption or operational integration.
The review's coding of 65 scholarly studies found extensive discussion of upskilling, reskilling, lifelong learning, partnerships, modular curricula, digital delivery, quality assurance, faculty incentives, and funding, but limited adoption measurement and operational integration evidence.
At the level of ordinary firms using generative AI, access is cheap and broadly equal even though the providers of the underlying models may possess concentrated market power.
Conceptual distinction between the foundation-model supply layer and the application/use layer, supported by the paper's literature synthesis.
Human calibration can improve local accuracy but limits plug-and-play scalability because companies or sectors may require bespoke prompt anchoring.
The paper's implications connect the observed benefits of company-contextualized prompting with the need for customer-specific calibration and reusable anchor examples.
Frontier AI development occurs mainly in advanced economies, but the implications for developing countries (the Global South) are equally significant and interconnected.
Paper asserts this as context and motivation; claim is based on literature synthesis and conceptual reasoning rather than reported empirical sample or specific data in the provided text.
The paper examines applications of generative AI across customer service, marketing, software development, healthcare, finance, law, logistics, and the creative industries.
Scope and domain-by-domain review described in the paper.
A structured cross-regional survey of tourism organizations reports widespread awareness and adoption of data/ML alongside uneven technological readiness.
Primary data from a structured cross-regional survey of tourism organizations described in the paper (survey method; specific sample size and regions not stated in the abstract).
There is notable variation across different AI coding agents in reviewer interaction metrics and merge rates.
Comparative analysis of reviewer interaction metrics and merge rates across five AI coding agents using the AIDev dataset.
Pull request description styles are associated with differences in merge outcomes.
Empirical examination of associations between PR description characteristics and merge outcomes (merge rates) across PRs from five AI coding agents in the AIDev dataset.
AI is currently adopted most comfortably in low-risk assistive uses—particularly summarization, document clarification, and preliminary review of lengthy narratives—rather than as a stand-alone engine for core accounting decisions.
Findings from thematic analysis of semi-structured interviews with 45 accounting practitioners across multiple sectors, as reported in the paper.