Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
This is the first study to quantify how much unsupervised algorithms accelerate the labeling process, and the first to compare labeling time from scratch to labeling time when using unsupervised algorithms as a pre-annotation step.
Novelty claim stated by the authors in the paper (literature positioning / authors' assertion).
Using unsupervised computer vision algorithms, the time required for the labeling process can be reduced from 170 hours to 37 hours, achieving an approximate reduction of 78%.
Empirical measurement reported in the paper comparing total labeling time when labeling from scratch (170 hours) versus using unsupervised algorithms as a pre-annotation step (37 hours).
Expert operators maintained a verification loop by persistently scanning the environment even when using LLM guidance.
Eye-tracking and behavioral data from expert participants showing continued environmental scanning (fixation metrics) and cross-referencing behavior in LLM-guided conditions.
LLM guidance enhanced task efficiency (higher rewards and victims-per-step) relative to a no-LLM baseline.
Experimental comparison in a simulated search-and-rescue environment across two LLM-guided conditions and a no-LLM baseline; behavioral measures reported for rewards and victims-per-step (eye-tracking and planning behavior also collected).
The contribution is a forecasting model and managerial planning tool for the shift to AI-augmented talent ROI accounting.
Theoretical and methodological development described in the paper (model + managerial guidance).
Output-based firms are forecast to outperform time-based peers by 1.5-2.0 percentage points in firm-level TFP growth by 2032.
Forecast produced by the paper's forecasting model for the transition to output-based talent accounting (modeling/forecasting exercise, not a realized empirical estimate).
Callaway-Sant'Anna doubly-robust staggered DiD estimates show a +4.51 percentage point increase in SG&A-to-revenue at t = +4, further supporting a positive overhead-pressure effect.
Callaway-Sant'Anna staggered DiD (doubly-robust) applied to the panel; point estimate at t = +4 reported.
Pooled event-study estimates show a +4.21 percentage point increase in SG&A-to-revenue at t = +3 (p = 0.001), consistent with an overhead-pressure signature.
Pooled event-study analysis on the panel; point estimate and p-value reported in the paper.
Under the revenue-percentile cohort proxy, a two-way fixed effects estimate shows an increase of +1.56 percentage points in SG&A-to-revenue (p = 0.049), indicating positive overhead pressure.
Two-way fixed effects regression on the DART panel using revenue-percentile cohort proxy; statistical significance reported (p = 0.049).
In a DART panel of 365 listed firms (2,281 firm-year observations), the SG&A-to-revenue ratio rose from 18.26 percent in 2018 to 20.06 percent in 2020, corrected mildly in 2021-2022, and peaked at 20.10 percent in 2024.
Descriptive statistics from the DART panel (365 firms; 2,281 firm-year observations) reported in the paper.
Korea's staged 52-hour workweek mandate provides an empirical early-warning case for overhead-pressure in the pre-τ regime.
Empirical analysis using a DART panel of listed Korean firms described in the paper.
The paper develops a forecasting framework for the transition from time-based talent accounting to output-based talent ROI in the human-AI era, centred on Theorem 3 (ROI Inversion at τ*).
Presentation of a theoretical forecasting model and formal theorem (Theorem 3) in the paper; methodological contribution rather than empirical test.
The paper derives C‑first policy prescriptions and offers three empirically testable propositions along with a falsifiable 10-year forecast.
Policy recommendations and empirical propositions presented in the paper (theoretical/policy-design evidence; forecast statement).
Convergence capacity (C) is distinct from absorptive capacity, dynamic capability, and human capital, and constitutes the specific cognitive mediator prior frameworks have left implicit.
Conceptual/definitional analysis and differentiation provided in the paper; theoretical argument distinguishing constructs.
A descriptive cross-national analysis of 20 OECD economies shows the AI × C interaction is associated with 86% of TFP variance, versus 31% for AI alone.
Empirical descriptive cross-national analysis reported in the paper; sample explicitly stated as 20 OECD economies (small-n analysis).
Using H-hat in the production function Y = F(K, H-hat) provides a human-centered mechanism for Solow's TFP residual: A_Solow = [1 + phi(A,C)]^(1-alpha).
Algebraic derivation connecting the proposed ICH augmentation factor to the Solow TFP residual (presented as a theoretical result in the paper).
The paper proposes the Intellectually Converged Human (ICH) framework with H-hat = H[1 + phi(A,C)], where effective productive capacity equals human capital (H) scaled by augmentation factor [1 + phi], and phi is jointly determined by AI utilization intensity (A) and convergence capacity (C).
Formal theoretical/model proposal presented in the paper (algebraic expression defining H-hat).
Digital infrastructure investment (computing power/NSC deployment) can be used as a policy instrument to correct excessive corporate financialization and guide corporate resources back to the real economy.
Interpretation and policy implication drawn from empirical results showing reduced financialization and increased real investment following NSC deployment.
Computing power deployment raises capital expenditure intensity.
Extended analysis of capital expenditure intensity metrics at the firm level following NSC establishment.
Computing power deployment increases firms' R&D investment.
Extended analysis using firm-level R&D spending in the post-NSC-deployment period.
Computing power deployment promotes reallocation to real investment, with significant increases in fixed assets investment.
Extended analysis of firm investment outcomes after NSC deployment (firm-level fixed-asset investment indicators).
Computing power deployment improves intelligent decision-making efficiency within firms, which increases core business returns and weakens incentives to hold financial assets.
Mechanism analysis in the paper using firm-level indicators of decision efficiency and performance, exploiting NSC staggered deployment.
Computing power deployment enhances firms' data-factor capitalization capability, which helps strengthen core business returns and reduces the motivation to allocate funds to financial assets.
Mechanism analysis in the empirical study (mediation/empirical channel tests) using firm-level data from Chinese A-share listed companies and variation from NSC rollouts.
Digital adoption has 56.6% larger impacts in high-standard markets (heterogeneity result).
Heterogeneity analysis reported in Results using Callaway & Sant'Anna estimator on the 8,547-firm panel; reported percentage larger impact in high-standard markets.
High-risk products show 84.1% stronger effects of digital adoption (heterogeneity result).
Heterogeneity analysis reported in Results on the same panel and estimation method; reported percentage stronger effect for high-risk products.
SMEs benefit 70.8% more from digital technology adoption than large firms (heterogeneity result).
Heterogeneity analysis reported in Results using the panel (8,547 firms) and staggered DiD estimator; reported percentage difference between SMEs and large firms.
Digital technology adoption promotes certification acquisition (mechanism test: coefficient = 0.286, p < 0.001).
Mechanism testing reported in Results using the same panel and estimation strategy; reported regression coefficient and p-value linking digital adoption to certification acquisition.
Effects of digital adoption intensify over time: long-term impacts reach 33.8%, which is 2.3 times the short-term effects.
Dynamic analysis reported in Results based on the same panel and staggered DiD approach; reported long-term percentage and multiplier relative to short-term effect.
Digital technology adoption increases certification acquisition by 55.0%.
Same panel (8,547 firms, 42 countries, 2015–2023) using staggered DiD estimator; reported point estimate in Results.
Digital technology adoption expands product scope by 14.9%.
Same panel (8,547 firms, 42 countries, 2015–2023) analyzed with Callaway & Sant'Anna estimator; reported point estimate in Results.
Digital technology adoption raises market entry probability by 12.8 percentage points.
Same panel (8,547 firms, 42 countries, 2015–2023) using Callaway & Sant'Anna staggered DiD estimator; reported point estimate in Results.
Digital technology adoption increases export value by 23.7%.
Panel data of 8,547 large-scale agricultural exporting firms from 42 developing countries (2015–2023) analyzed with the Callaway and Sant'Anna staggered difference-in-differences estimator; reported point estimate in Results section.
The paper concludes with specific policy recommendations addressing procurement, workforce development, standards alignment, and interagency coordination to accelerate responsible AI adoption across the federal audit ecosystem.
Statement of the paper's conclusions and policy recommendations (descriptive of paper content). No empirical evaluation reported for the effectiveness of these recommendations.
Critical success factors for AI-augmented audit include executive sponsorship at the agency leadership level, dedicated cross-functional implementation teams with embedded data science competencies, iterative pilot deployments that generate performance evidence prior to enterprise rollout, and robust governance structures that maintain human judgment at consequential decision points.
Paper's recommended critical success factors based on synthesis of implementations and best-practice guidance; presented as prescriptive guidance rather than validated causal evidence.
A structured three-phase implementation approach spanning 24 to 48 months enables federal audit agencies to achieve meaningful AI augmentation of core audit functions while managing implementation risk within acceptable bounds.
Paper's proposed implementation timeline and argument (recommendation based on the paper's synthesis). No empirical test or sample size reported to validate the timeline.
The paper draws on recent advances in intelligent fraud monitoring, machine identity governance, adaptive risk scoring, and digital forensics analytics to ground its recommendations in the most current available evidence on AI audit capability development.
Paper cites and synthesizes recent technical advances and implementations in specific AI audit subdomains (literature/implementation synthesis). No sample sizes or systematic review metrics provided.
The roadmap addresses four core implementation domains: technical infrastructure and data architecture requirements; human capital and organizational change management for audit workforce transformation; governance, ethics, and risk management frameworks; and policy and standards development to enable AI-augmented oversight.
Paper's stated structure and recommendations (categorization of implementation domains). Descriptive; no quantitative evaluation reported.
The paper develops an original conceptual framework designated the AI-Augmented Audit Continuum (AIAC) to guide progressive capability development from foundational analytics to autonomous audit functions.
Paper claims and framework development (conceptual contribution). No empirical validation or sample size reported.
This paper develops a comprehensive policy and implementation roadmap for the deployment of AI-augmented audit capabilities within United States government agencies and multilateral organizations, synthesizing evidence and aligning strategies with GAO, OMB, and INTOSAI frameworks.
Statement of the paper's scope and methods (synthesis of evidence; alignment analysis with GAO, OMB, INTOSAI). This is a description of the paper's contribution rather than an empirical finding.
Artificial intelligence technologies, including machine learning, natural language processing, network analytics, and intelligent process automation, offer substantial potential to augment the analytical capacity of public audit institutions, extend audit coverage to previously inaccessible transaction populations, and accelerate detection timelines from years to days or hours.
Author's synthesis and claims in the paper; references to existing AI audit implementations across federal, state, and international contexts (literature/implementation synthesis). No specific sample size reported.
Utility-Aligned Profile Exploration generates multiple candidate profiles per cluster, evaluates them via a lightweight downstream utility proxy, iteratively refines the best candidates and constructs preference pairs for DPO fine-tuning.
Methodological description in the paper of the profiling and DPO fine-tuning pipeline; empirical benefit supported elsewhere in the paper by reported metrics.
Tool-Augmented Global Knowledge Mining equips an LLM agent with 27 analytical tools to mine platform-scale data, producing reusable global knowledge, adaptive user clustering rules, and region-level supply-demand priors.
Methodological description in the paper reporting 27 analytical tools and the outputs produced by that module.
ProfiLLM was deployed on DiDi's production dispatcher.
Stated deployment in the paper; deployment is described as production integration (no further deployment metrics in the abstract).
In a 14-day online A/B test, ProfiLLM produced consistent improvements including -0.82% Cancel-Before-Accept rate.
14-day live online A/B experiment on DiDi's production dispatcher; precise sample size/statistical significance not stated in the abstract.
In a 14-day online A/B test, ProfiLLM produced consistent improvements including +0.33% Completion Rate.
14-day live online A/B experiment on DiDi's production dispatcher; sample counts not provided in the abstract.
In a 14-day online A/B test, ProfiLLM produced consistent improvements including +0.47% GMV.
14-day live online A/B experiment on DiDi's production dispatcher; number of users/orders not specified in the abstract.
ProfiLLM achieves up to +4.35% GMV gain in dispatching simulation.
Dispatching simulation results reported in paper; specifics of simulation scale/replicates not provided in the abstract.
ProfiLLM achieves up to +6.14% relative AUC improvement in outcome prediction.
Quantitative evaluation reported in paper after deploying ProfiLLM on DiDi's production dispatcher; exact test set/sample not stated in the abstract.
From a practical perspective, the study offers a conceptual measurement framework and policy guidance for municipal decision makers seeking to improve productivity while strengthening resilience and reducing systemic risks in increasingly interconnected public governance systems.
Paper presents a conceptual measurement framework and policy recommendations derived from the integrative review and framework; asserted in discussion and implications sections.
Resilience depends on the ability of public organisations to anticipate, absorb, adapt to, and recover from AI-related disruptions while maintaining the continuity and quality of public services.
Theoretical framing (sociotechnical systems and resilience theory) supported by synthesis of reviewed empirical studies; proposed conceptual measurement framework in the paper.