Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
PRISM automatically generates test cases from plain-language agent requirements.
Methodological description in abstract stating PRISM takes plain-language requirements and automatically generates test cases.
PRISM successfully identifies and repairs production regressions caused by LLM behavioral drift within a 24-hour detection window.
Reported result in abstract claiming detection and repair of production regressions within 24 hours during deployment.
PRISM achieves 99% production reliability across all evaluated agents.
Reported quantitative outcome in abstract from the three-week evaluation over 35 agents.
PRISM reduces median prompt authoring time from 2 days to under 30 minutes.
Reported quantitative outcome in abstract from the three-week evaluation across the 35 agents.
Policymakers should combine support for technological development with strategic investments in finance, trade integration, and public infrastructure to maximize AI's economic benefits and transform its potential into sustainable and inclusive growth.
Policy recommendation derived from the empirical findings (positive AI effects and positive interactions with financial innovation, trade openness, and government consumption) reported for 19 G20 countries (2005–2023) using GMM.
The interaction between AI and government final consumption expenditure helps strengthen economic growth by improving public infrastructure, institutional quality, and capacity to leverage new technologies.
GMM interaction specifications using panel data for 19 G20 countries (2005–2023); reported AI × government final consumption expenditure interaction coefficient is positive and statistically significant, with interpretation linking it to public infrastructure and institutional capacity.
The interaction between AI and trade openness is positive and significant, underscoring the role of international trade in technological diffusion and competitiveness to boost growth.
GMM interaction models on panel data (19 G20 countries, 2005–2023); reported AI × trade openness interaction coefficient is positive and statistically significant.
The interaction between AI and financial innovation has a positive and significant impact on economic growth, indicating that innovative finance mediates AI's technological potential into tangible economic gains.
GMM models with interaction terms using panel data of 19 G20 countries (2005–2023); reported AI × financial innovation interaction coefficient is positive and statistically significant.
AI-related innovation has a positive and significant effect on economic growth (linear model, GMM).
Panel analysis of 19 G20 countries (2005–2023) using the Generalized Method of Moments (GMM) linear model; reported positive and statistically significant coefficient for AI-related innovation.
Synthetic scenarios in the paper illustrate that the revised metric distinguishes between frequent low-leverage use, semantically repetitive prompting, and more autonomous, higher-consequence AI-assisted work.
Paper includes synthetic scenario simulations/illustrations demonstrating metric behavior across different usage patterns (synthetic examples; no real-world sample reported).
The authors derive sub-daily update rules and a bounded interpretation layer for estimated efficiency and financial impact from the IIQ metric.
Analytic derivation in the methods: paper presents update rules (sub-daily) and an interpretation layer mapping IIQ to estimated efficiency and financial impact (theoretical derivation / worked examples). No empirical validation sample reported.
The formulation produces a raw Intelligence Adoption Index (IAI) and a normalized 0-1000 IIQ index for comparison between heterogeneous users and units.
Methodological description: authors define a raw IAI and describe normalization to a 0–1000 IIQ scale for comparability (model/specification). No empirical sample reported.
IIQ combines a novelty-weighted, time-decayed token stock with usage frequency, a grace-period recency gate, organizational leverage, task complexity, and autonomy to form its measurement.
Methodological formulation in the paper: component-level specification of the IIQ metric (theoretical specification / algorithmic description). No empirical validation sample reported.
The Intelligence Impact Quotient (IIQ) is a composite metric intended to quantify the depth to which AI systems are integrated into organizational work and their impact.
Paper framing and definition: the authors introduce IIQ as a composite metric and describe its purpose as quantifying AI integration depth and impact (conceptual/methodological description). No empirical sample reported.
Das Dokument leistet einen Beitrag zu den laufenden Bemühungen der G7 und der OECD, die Verbreitung innovativer, vertrauenswürdiger und produktivitätssteigernder KI im Einklang mit den KI-Grundsätzen der OECD zu fördern.
Descriptive claim about the paper's intended contribution to G7/OECD efforts and alignment with OECD AI Principles (self-declared by the paper).
Die Erkenntnisse unterstreichen, dass die Regierungen Strategien unterstützen sollten, die die Einführung von KI in KMU beschleunigen und eine digitale Transformation fördern, die allen zugutekommt.
Policy recommendation based on the paper's synthesis of data analysis and case studies; presented as the paper's conclusion (no causal estimate provided in excerpt).
Das Dokument führt eine Taxonomie der KI-nutzenden KMU auf Basis des digitalen Reifegrads, der Komplexität der Nutzung und des Umfangs der Anwendung ein, die darauf abzielt, die Politikgestaltung zu unterstützen.
Descriptive statement that the paper develops a taxonomy (method: taxonomy construction based on those three dimensions); presented as part of the paper's contributions (no empirical validation details in excerpt).
Dieses Diskussionspapier wurde auf Ersuchen der G7-Präsidentschaft vom OECD-Sekretariat erstellt, um Hintergrundmaterial für die Diskussionen der G7 über einen Blueprint zur Einführung von KI in KMU bereitzustellen.
Descriptive statement of the paper's provenance and purpose (administrative/factual about document preparation).
Im Rahmen der G7-Präsidentschaft Kanadas 2025 wurde die beschleunigte Einführung von KI in KMU zu einer Hauptpriorität erklärt.
Factual statement about G7 policy priorities as reported by the paper (administrative/policy fact reported by OECD secretariat).
Künstliche Intelligenz (KI) ist ein vielversprechender Ansatz, um Produktivität und Innovation in Unternehmen, insbesondere kleinen und mittleren Unternehmen (KMU), zu steigern.
Authoritative statement in the paper's summary; based on literature review and general argumentation rather than a specific empirical test reported in this excerpt.
The research contributes to the literature on technology adoption in developing economies and offers policymakers and business leaders in sub-Saharan Africa valuable insights.
Paper's stated contribution in the abstract; a general claim about the study's scholarly and policy relevance rather than a quantifiable empirical result.
Targeted policy interventions — such as upskilling initiatives and supportive regulatory frameworks — are important to harness AI’s benefits while mitigating adverse impacts on workers.
Paper conclusion/recommendation drawn from empirical findings (positive association of AI with productivity and sales, plus observed cross-country variation). This is presented as a policy implication; no empirical evaluation of specific policies is reported in the excerpt.
AI adoption has a significant positive relationship with firm sales growth in the selected sub-Saharan African countries.
Same firm-level World Bank Enterprise Surveys (2007–2024) and regression methods (FGLS, robust OLS, HDFE) as above. Paper statement: "AI has a significant positive relationship with ... sales growth." Exact sample size and numeric effect not provided in excerpt.
AI adoption has a significant positive relationship with firm labour productivity in the selected sub-Saharan African countries.
Firm-level dataset from the World Bank Enterprise Surveys covering 2007–2024; empirical analysis using feasible generalized least squares (FGLS), robust OLS, and high-dimensional fixed effects (HDFE) linear regressions. Paper statement: "AI has a significant positive relationship with firm labour productivity." Exact firm sample size not reported in the provided excerpt.
There is a positive spillover effect on AI-ineligible chats: treated workers adapted their multitasking workflow to devote greater attention to these chats.
Experiment-level observations comparing worker behavior on AI-ineligible chats between treatment and control; treated workers reallocated attention/effort (multitasking workflow changes) leading to improved attention on AI-ineligible chats.
Early intervention is essential for sustaining high post-escalation intervention effort.
Temporal analysis of intervention timing within the randomized experiment showing an association between earlier human intervention after escalation and higher subsequent intervention effort.
Human intervention preserves service quality in algorithm-triggered technical escalations (unresolved customer issues beyond the AI's capability).
Experimental subgroup analysis of escalations categorized as algorithm-triggered technical escalations; post-escalation human interventions were observed to maintain service quality in these cases.
PRIF yielded an average ROI of 83%.
Reported financial evaluation/ROI estimate following PRIF adoption in the paper (derived from pilot/case study cost-benefit or sample analysis).
PRIF adoption reduced financial misstatements by 47%.
Reported change in financial misstatement incidence after PRIF implementation in the paper's evaluation (case studies/forensic report analysis).
PRIF adoption reduced compliance resolution time by 58%.
Reported performance metric after PRIF adoption in pilot/case studies described in the paper.
Client retention was 91% for high SCI versus 54% for low SCI.
Reported retention rates stratified by SCI levels in paper (presumably derived from the sample used for SCI analysis).
The Stakeholder Communication Index (SCI) revealed a strong correlation (r = 0.83) between report quality and client retention.
Statistical analysis reported in paper linking SCI-derived report quality scores to client retention; correlation coefficient r = 0.83 provided.
Accuracy increased from 62% to 89–94% after integration of AI and blockchain.
Reported accuracy figures in results section based on PRIF evaluation (presumably from analyzed forensic reports/case studies).
Integration of AI and blockchain reduced the risk detection time from 47 days post-event to 9–22 days pre-event.
Reported results from PRIF implementation/pilot using case studies and forensic report analysis (paper cites these temporal comparisons).
This study pioneers a Proactive Risk Intelligence Framework (PRIF) for Chartered Accountant (CA) firms, targeting gaps in risk anticipation, stakeholder communication, and compliance.
Paper description of study objective and framework development (mixed-method design, interviews, case studies, forensic report analysis).
We distill practical design principles for selecting logging policies when operational constraints prevent implementing the theoretical optimum.
Abstract states the paper provides practical design principles derived from the theoretical work; basis is methodological/theoretical synthesis (no empirical sample size provided in abstract).
We demonstrate the importance of treatment selection when gathering data for OPE, and describe theoretically optimal approaches when this is a firm's primary objective.
Abstract claims demonstration and description of theoretically optimal approaches; evidence likely consists of analytical results and/or illustrative demonstrations in the paper (no sample size reported in abstract).
We propose a unifying framework for logging policy design and derive optimal policies in canonical informational regimes where the target policy and reward distribution are (i) known, (ii) unknown, and (iii) partially known through priors or noisy estimates at logging time.
Abstract states the development of a framework and derivations of optimal policies across specified informational regimes; evidence is theoretical derivations (no empirical details in abstract).
In practice OPE accuracy depends heavily on the logging policy used to collect data for computing the estimate.
Assertion in abstract motivated by the authors' study of logging policy design; implies analytical results in the paper relating logging policy to OPE error (no sample size given in abstract).
Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy, enabling high-stakes experimentation without live deployment.
Statement in abstract describing OPE and its role; conceptual/theoretical description (no sample size or empirical study reported in the abstract).
These verified assertions improve users' performance on code-comprehension tasks in a user study with more than 400 participants.
User study reported in the paper: a study involving more than 400 participants measured performance on code-comprehension tasks with and without the verified assertions (sample size reported as >400 participants).
Evaluation on 18 diverse programming tasks suggests that Viverra can efficiently generate code with verified assertions.
Empirical evaluation reported in the paper: a test set of 18 programming tasks was used to evaluate Viverra's ability to generate code with verified assertions (sample size = 18 tasks).
Viverra verifies those assertions in a compositional and best-effort manner via a portfolio of bounded model checkers.
Method description: the paper states that verification is done compositionally and in a best-effort way using a portfolio of bounded model checkers (implementation/algorithmic claim).
Given a natural-language task description, Viverra prompts an LLM to synthesize a C program together with candidate assertions expressing safety and correctness properties.
Method section description: the workflow described in the paper explicitly states LLM prompting to produce C programs and candidate assertions (methodological claim, illustrated with examples).
Viverra automatically produces formally verified annotations alongside generated code to aid users' understanding of the generated program.
System description in the paper: Viverra is presented as a system that generates code together with formally verified annotations; implementation details and demonstration are described (no precise external benchmark cited here).
Adversarial inputs evolved using a small proxy model retain high effectiveness against large commercial LRMs (strong transferability).
Reported transfer experiments in the abstract showing that evolved adversarial inputs from a small proxy model remain effective against larger commercial models; no numeric transfer success rates provided in abstract.
Across four state-of-the-art reasoning models, the proposed method substantially amplifies output length, achieving up to a 26.1x increase on the MATH benchmark and consistently outperforming benign and manually crafted missing-premise baselines.
Experimental results reported in the abstract: evaluations on the MATH benchmark and comparisons against benign and manually crafted missing-premise baselines across four SOTA models.
We propose an automated black-box framework that induces overthinking in LRMs by systematically perturbing the logical structure of input problems using a hierarchical genetic algorithm (HGA) operating on structured problem decompositions and optimizing a composite fitness function to maximize response length and reflective overthinking markers.
Methodological description of the proposed approach (HGA and composite fitness) as presented by the authors in the abstract.
Function signatures, constraints and style descriptions emerge as the most influential prompt dimensions affecting the readability of LLM-generated code.
Systematic examination of multiple prompt dimensions in the paper, reporting that function signatures, constraints, and style descriptions had the largest measured influence on readability scores.
We evaluate the readability of code generated by mainstream LLMs under 5,869 scenarios extracted from large code bases including World of Code (WoC) and LeetCode.
Empirical evaluation reported in paper using 5,869 scenarios drawn from WoC and LeetCode; LLM-generated code samples were produced and scored with the readability model.