Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
In that deployment the framework measured approximately $378,000 in annual labor value of machine-equivalent work.
Same empirical manufacturing deployment reported in the paper (single case/example).
In a representative manufacturing deployment, the framework measured 8.4 FTE of machine-equivalent labor.
Empirical example reported in the paper described as a 'representative manufacturing deployment' (appears to be a single deployment/case).
The paper introduces the Machine Labor Index (HEWU-PSI), a time-series economic indicator designed to track aggregate machine labor output at company, sector, and national level, analogous in function to the Purchasing Managers' Index.
Methodological contribution described in the paper (proposal of an index and its intended scope; no empirical time-series dataset reported).
The paper introduces AILU (AI Labor Units) as a software-specific subset metric.
Methodological contribution described in the paper (definition of a software-specific metric subset).
The paper presents the conceptual foundation, mathematical model (HEWU = MO ÷ HB × CF × QF), calibration framework, Baseline Library architecture, and auditability mechanisms underlying the standard.
Paper's methodological content (explicit model formula and supporting frameworks described).
This paper introduces the Human-Equivalent Work Unit (HEWU), a standardized metric that converts AI and automation system output into human labor equivalents, expressed as full-time employee (FTE) equivalents and annual labor value ($).
Methodological contribution described in the paper (definition and proposal of a new metric; no empirical validation sample reported).
Artificial intelligence systems are autonomous agents performing economically meaningful labor at scale across customer service, software engineering, logistics, manufacturing, and knowledge work.
Author's conceptual/empirical assertion in the paper (no specific sample, presented as general observation).
Our findings indicate an increasing agent activity in open-source projects.
Trend analysis reported in the paper showing growth in agent-originated activity within the assembled dataset of PRs and associated metadata.
TAI introduces recursive feedback loops between technology, knowledge, and output that redefine long-term growth trajectories and the equilibrium conditions of economies.
Derived from the paper's dynamic model: analytical results showing feedback mechanisms between technology, knowledge stock, and output; presented as theoretical model implications rather than validated empirical findings.
The model integrates AI as both a productivity amplifier and an autonomous driver of capital accumulation.
Stated methodological contribution: the authors extend Solow (1956) and Romer (1990) frameworks to build a dynamic model in which AI enters production as an amplifier of productivity and as an autonomous engine for capital accumulation; evidence is theoretical/model construction rather than empirical.
Transformative artificial intelligence (TAI) is capable of driving structural economic change comparable to the industrial revolution.
The paper asserts this claim by analogy and conceptual argument in the introduction; it frames TAI as 'capable of driving structural economic change comparable to the industrial revolution' without reporting empirical data — supported by theoretical reasoning and historical analogy.
Implicit budget constraints from BCR circumvent adversarial gradients and catastrophic optimization collapse that occur with explicit length penalties, providing a highly stable, constraint-based alternative for length control.
Empirical comparison between BCR (implicit budget constraints) and methods using explicit length penalties reported in the paper; claim of improved stability and avoidance of catastrophic optimization collapse.
Qualitative analyses reveal emergent self-regulated efficiency: models autonomously eliminate redundant metacognitive loops without explicit length supervision.
Qualitative analysis of model behavior reported in the paper (no quantitative effect sizes provided in the excerpt).
BCR challenges the traditional accuracy-efficiency trade-off by demonstrating a 'free lunch' phenomenon at standard single-problem inference (i.e., reduced token usage with maintained or improved accuracy even at N=1).
Reported experimental results on 1.5B and 4B model families showing token reductions and maintained/improved accuracy at standard single-problem inference.
As N increases, accuracy degrades far more gracefully than baselines, establishing N as a controllable throughput dimension.
Comparative experiments versus baselines varying concurrent-problem count N; qualitative claim that accuracy degradation is 'far more graceful' than baselines.
As the number of concurrent problems N increases during inference, per-problem token usage decreases monotonically.
Reported experimental finding described as a novel task-scaling law observed when varying N at inference time; no numeric effect sizes provided in the excerpt.
Batched Contextual Reinforcement (BCR) reduces token usage by 15.8% to 62.6% while consistently maintaining or improving accuracy across five major mathematical benchmarks.
Empirical evaluation reported in the paper across two model families (1.5B and 4B) and five mathematical benchmarks; token usage reduction range and qualitative accuracy statement provided.
Results may be applied in the development of financial institution strategies, regulatory frameworks, risk management systems and professional training programmes.
Applied implications drawn from the literature synthesis and comparative analysis; presented as potential uses rather than empirically validated interventions.
Significant changes in human resource needs are occurring, with growing demand for analysts and specialists combining financial and technological competencies.
Conclusion from literature review and synthesis of international studies on labour demand in finance under Big Data/AI adoption; no original labour-market survey included.
Big Data and AI technologies significantly improve efficiency, risk assessment accuracy, fraud detection and financial inclusion.
The paper reports results from a qualitative analysis of recent academic literature, comparative analysis of sector-specific applications, and synthesis of empirical findings from international studies; no primary sample size reported.
Overall, findings highlight that AI serves as a revolutionary (transformative) tool rather than merely a replacement tool for employment—changing the nature of human work rather than simply disengaging it.
Synthesis conclusion in the paper drawing on the literature review and the authors' empirical results indicating task reallocation and changing job content.
The paper argues for equal technology governance as a necessary policy response to AI's labor market effects.
Policy recommendations discussed in the paper that call for equitable governance of AI; based on literature synthesis and empirical findings.
The analysis raises policy implications emphasizing reskilling and education to address AI-driven changes in the labor market.
Policy discussion section summarized in the paper; draws on empirical findings and literature to recommend reskilling/education.
Moderate AI usage is associated with employment growth.
Part of the U-shaped relationship reported in the paper's empirical results; described qualitatively in the abstract/summary.
Secondary empirical evidence from Colombia's EDIT manufacturing survey (N=6,799 firms) shows that management practice quality amplifies the return to technology investment (interaction coefficient 0.304, p<0.01).
Secondary empirical analysis of EDIT manufacturing survey data; sample size reported as N = 6,799 firms; regression interaction term reported as coefficient 0.304 with p < 0.01.
We endogenize the augmentation function as phi(D, W), where W is a five-dimensional workplace design vector (AI interface design, decision authority allocation, task orchestration, learning loop architecture, psychosocial work environment), and prove that human-centric design is profit-maximizing when the workforce's augmentable cognitive capital exceeds a critical threshold.
Theoretical model and formal proof presented in the paper (analytical derivation of phi(D,W) and threshold condition).
There is a need for energy-efficient AI development to align technological progress with sustainable energy consumption.
Policy recommendation based on the paper's empirical findings that AI adoption increases firm-level electricity demands in the short run; normative argument rather than a directly tested empirical claim.
The AI-related widening of the electricity output growth gap is stronger among manufacturing firms, non-state-owned firms, small firms, low-tech firms, and low-energy-consumption and low-pollution firms.
Heterogeneity/subgroup analyses across firm characteristics (ownership type, size, sector, technology intensity, baseline energy use and pollution levels) showing larger estimated effects in the listed subgroups. Specific subgroup sample sizes and coefficients not reported in the summary.
The effect of AI adoption on the electricity output growth gap is more pronounced for firms operating in highly competitive industries.
Heterogeneity analysis by industry competition intensity (likely via industry-level measures of competition); interaction regressions showing larger estimated effects in more competitive sectors. Sample/subgroup sizes not specified in the summary.
The effect of AI adoption on widening the electricity output growth gap is more pronounced for firms located in economically advanced regions.
Heterogeneity analysis by regional economic development level using the firm-level electricity consumption dataset; stratified or interaction regressions showing larger estimated effects in more advanced regions. Exact subgroup sizes not provided in the summary.
The main result (initial widening of electricity growth gap) is robust to alternative variable definitions, exclusion of firms relying on outsourced AI services or non-AI adoption samples, and controls for endogeneity.
Robustness checks reported in the paper: alternative variable definitions, sample restrictions (excluding outsourced-AI-reliant firms and non-AI samples), and application of endogeneity control methods (e.g., instrumental variables or panel fixed effects). Exact methods and sample sizes not specified in the summary.
AI adoption initially widens the corporate electricity output growth gap at the firm level in China.
Empirical analysis using unique firm-level data on corporate electricity consumption in China; econometric estimation comparing electricity output growth between AI-adopting firms and non-adopting peers (panel/firm-level analysis). Sample size not stated in the summary.
To optimize agentic AI integration and ensure responsible innovation across financial services, interdisciplinary, longitudinal research and robust governance frameworks are needed.
Authors' conclusions and recommendations based on the identified findings and gaps in the reviewed literature.
Diverse architectural models such as multi-agent systems and cloud-based frameworks enable scalable, adaptive agentic AI deployments in financial services.
Synthesis of architecture-focused studies and framework descriptions within the reviewed literature (architectural benchmarking across papers).
Findings reveal substantial productivity gains and operational efficiencies predominantly in banking and investment.
Systematic review synthesizing multidisciplinary qualitative, quantitative, and bibliometric studies of agentic AI applications in financial services published up to mid-2024 (review-level synthesis).
The ManagerWorker two-agent pipeline (expensive text-only manager + cheaper worker with repo access) can substitute expensive execution by using expensive reasoning in the manager and cheaper execution in the worker.
System design description plus empirical results on 200 SWE-bench Lite instances showing parity in success rates between a strong-manager/weak-worker pipeline and a strong single agent while using fewer strong-model tokens.
A minimal review-only manager loop adds only 2 percentage points over the baseline, whereas structured exploration and planning by the manager add 11 percentage points, demonstrating that active direction (not mere reviewing) produces most of the benefit.
Ablation-style comparison of pipeline variants on the 200-instance SWE-bench Lite evaluation: review-only manager loop versus manager with structured exploration and planning; reported improvements in percentage points.
A strong manager directing a weak worker achieves a 62% success rate on software-engineering tasks, matching a strong single agent which achieves 60%, while using a fraction of the strong-model token usage.
Empirical evaluation on 200 instances from SWE-bench Lite across five pipeline configurations and model pairings; measured task success rates and token usage for manager-worker pipelines versus single-agent baselines.
Under economy-wide deployment, the share of computer-vision-exposed labor compensation that is cost-effectively automatable rises sharply (relative to the firm-level 11% estimate).
Model counterfactuals or calibration scenarios comparing firm-level deployment vs economy-wide deployment; qualitative statement that share increases substantially.
At the firm level, cost-effective automation captures approximately 11% of computer-vision-exposed labor compensation.
Calibration and implementation in computer vision; reported firm-level estimate from the framework.
Scale of deployment is a key determinant: AI-as-a-Service and AI agents spread fixed costs across users, sharply expanding economically viable tasks.
Modeling and calibration arguments showing fixed-cost spreading effects increase set of tasks for which automation is cost-effective; qualitative and quantitative comparisons in implementation.
Because higher accuracy is disproportionately costly (convex cost), full automation is often not cost-minimizing; partial automation, where firms retain human workers for residual tasks, frequently emerges as the equilibrium.
Theoretical model combined with calibration (scaling laws + task mappings); equilibrium outcomes reported from the framework implementation.
We model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, from no automation through partial human-AI collaboration to full automation.
The paper develops a theoretical framework / model that treats automation intensity as a continuous decision variable; described as the central modeling approach.
The findings demonstrate that technological innovation strategies, when effectively implemented, provide measurable competitive advantages for banks and offer evidence-based insights for policymakers and practitioners.
Authors' interpretation/conclusion drawing on the reported statistically significant relationships between innovation (product and technological) and competitiveness.
Technological innovation is positively and statistically significantly related to bank competitiveness (simple linear regression result reported).
Simple linear regression reported in the paper testing the hypothesis that technological innovation influences competitiveness; data collected from innovation-focused executives across licensed banks (paper states data from 39 licensed banks).
Product innovation strategy has a positive and statistically significant effect on competitiveness (F(1,134) = 74.983, p < .001).
Bivariate regression analysis reported in the paper with F(1,134)=74.983, p < .001; based on survey data from innovation-focused executives (regression degrees of freedom indicate n≈136 observations).
In the user study, AI-expanded 5W3H prompts increase user satisfaction from 3.16 to 4.04.
Reported pre/post or baseline vs AI-expanded satisfaction scores in the N=50 user study with numeric scores 3.16 and 4.04.
In the user study, AI-expanded 5W3H prompts reduce interaction rounds by 60 percent.
Reported comparison in the N=50 user study between baseline interaction rounds and rounds after AI-assisted 5W3H expansion; percentage reduction reported as 60%.
A weak-model compensation pattern was observed: the lowest-baseline model (Gemini) shows a much larger D-A gain (+1.006) than the strongest model (Claude, +0.217).
Model-level comparison of D-A gain (difference between structured and unstructured conditions) across three models (Claude, GPT-4o, Gemini) on the evaluated outputs; reported gains for Gemini and Claude.
The strongest structured conditions reduce cross-language sigma from 0.470 to about 0.020.
Reported numeric comparison of sigma (variance) between unstructured baseline and strongest structured prompting conditions across evaluated outputs.