Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The staggered expansion of Turkey's national natural gas pipeline network provides plausibly exogenous variation in connectivity because pipeline routing is determined by energy distribution priorities rather than digital demand.
Identification strategy described by the authors: using pipeline expansion as an instrument/conduit for fiber-optic deployment; argument rests on institutional routing rules and timing.
This is an exploratory and qualitative state-of-practice study grounded in over 30 interviews across four stakeholder groups (large enterprises, small/medium firms, AI developers, and CAD/CAM/CAE vendors).
Methodological statement in the paper describing study design and sample composition.
Key breakthroughs needed include integration with traditional engineering tools and data types, robust verification frameworks, and improved spatial and physical reasoning.
Interviewee-identified requirements compiled from over 30 interviews; stakeholders repeatedly pinpoint integration, verification, and spatial/physical reasoning as priority technical advances.
The goal is not to identify causal effects, but to document stylized facts about how technology changes the scale of asset management work.
Author's stated research objective in the paper's summary/introduction (explicitly notes descriptive, not causal, intent).
Using a small panel of representative firms, we compare changes in AUM per employee, revenue per employee, and operating expense intensity over time.
Stated empirical approach: analysis of a small panel of representative firms comparing three metrics (AUM/employee, revenue/employee, operating expense intensity) over time. The excerpt notes panel is 'small' but gives no numeric sample size or firm list.
This project studies how much labor is required to manage capital across those waves by tracking a simple productivity measure: assets under management per employee.
Stated research design: longitudinal tracking of assets under management (AUM) per employee as the primary productivity measure; described in the paper's methods/summary. No numeric sample size provided in the excerpt.
Financial firms have gone through three major technological waves: computerization in the 1980s and 1990s, the rise of indexing and passive investing in the 2000s and 2010s, and the AI and automation wave from roughly 2015 to the present.
Author's historical categorization stated in the paper's introduction/summary (time periods specified). No sample or empirical test reported in the excerpt.
The paper includes a companion video demonstrating the approach: https://youtu.be/55Q3lq1fINs.
Statement in paper providing link to companion video.
The physical robot scenario used a 7-DOF robot arm to validate the approach.
Experimental setup description in paper specifying hardware used (7-DOF robot arm).
Prior research typically considers task-level and motion-level adaptation in isolation (task-level methods ignore spatial interference; motion-level methods ignore broader task context).
Literature summary/related work section asserting the separation of prior task-level and motion-level approaches.
RAPIDDS models an individual's spatial behavior (motion paths) and temporal behavior (time required to complete tasks) over multiple cycles.
Description of modeling approach in the paper (method details describing spatial and temporal individual models over multi-cycle interactions).
This paper introduces RAPIDDS, a framework that unifies task-level and motion-level adaptation for human-robot teaming.
Methodological contribution described in paper (framework design and implementation).
This study proposes a framework for evaluating platform ecosystems by their long-term effects on human capital formation and institutional resilience.
Methodological contribution claimed by the paper (development of an evaluative framework); presented as part of the paper's contributions rather than an empirical finding.
Endogeneity in estimating AI's effects was controlled using a two-way fixed effects (TWFE) model and Propensity Score Matching (PSM).
Methodological claim reported in the study about the identification strategy used to estimate causal effects of AI adoption.
The timing of AI adoption was identified through a multi-step, contextually validated text analysis of DART business reports.
Descriptive/methodological statement in the study describing how adoption dates were extracted from firms' regulatory/business reports (DART) via a validated text-analytic procedure.
The average effect of AI adoption on market value (Tobin's Q) was not statistically significant across all firms.
TWFE and PSM estimates on KOSDAQ-listed firms (2018–2025) reporting firm-level Tobin's Q before and after identified AI-adoption timing.
No statistically significant change was observed in return on assets (ROA) following AI adoption.
Same empirical setting as above (KOSDAQ firms 2018–2025) using TWFE and PSM to estimate causal effects of AI adoption on ROA.
Our findings highlight the importance of additional research and progress on economic measurement related to AI.
Authors' concluding statement/recommendation based on their results and measurement challenges discussed in the paper.
We use tools to indirectly estimate the impact of AI via the lens of BEA’s industry accounts.
Methodological description in the paper: authors apply indirect estimation methods using BEA industry accounts to infer AI's economic impact.
Currently, there is not a line item in the U.S. national accounts that can be used to identify and measure the economic impact of artificial intelligence (AI).
Statement by authors about the state of U.S. national accounts (BEA) and absence of a specific national-accounts line item for AI.
We compare multiple state-of-the-art agents (e.g., GPT-4o, Llama 3, Qwen2) on metrics assessing tool selection accuracy, faithfulness, and hallucination.
Paper lists evaluated models (GPT-4o, Llama 3, Qwen2) and reports evaluation on metrics including tool selection accuracy, faithfulness, and hallucination across the benchmark.
Our benchmark consists of 100 financial questions.
Paper explicitly states the benchmark contains 100 financial questions.
Under three scenarios (optimistic: 2028-2035; base: 2035-2045; pessimistic: 2045-2060), we specify disconfirmation criteria that would weaken the thesis if observed.
Scenario analysis and specification of disconfirmation criteria by the authors; methodological claim about forecasting structure rather than empirical result.
Converging evidence from history, philosophy, neuroscience, technology, organizational studies, and cultural analysis supports this thesis.
Authors' multidisciplinary literature review and synthesis across the named fields (method: qualitative review); no single empirical dataset or sample size given.
We introduce 'instrumental dissolution' -- loss of institutional-default status while persisting in specialist niches.
Conceptual/theoretical contribution defined by the authors and illustrated via cross-disciplinary examples; no empirical validation sample reported.
Typing's dominance was instrumental, not cognitively necessary.
Argumentative/historical analysis presented in the paper; synthesis of historical and philosophical literature (no empirical sample or experiment reported).
We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews.
Reported evaluation methods and sample in the paper's abstract: log analysis, surveys, and 20 interviews with over 2,200 participants across 116 countries.
Participants were retested individually on the programming tasks after a retention interval of one week.
Statement in abstract describing follow-up retest procedure (one-week retention interval, individual retest).
Participants were incentivized by bonus compensation to balance performance with understanding.
Paper description of participant incentives in methods/abstract; compensation scheme used during experiment.
We conducted a controlled pair programming study with 22 participants who wrote Python code under time pressure in teams of two and individually with GitHub Copilot for 20 minutes each.
Statement of study design in the paper's methods/abstract; controlled pair programming experiment with 22 participants, 20-minute tasks in both conditions (human teammate and Copilot).
The framework is evaluated against forecast-driven base-stock and greedy fulfillment heuristics, and against a perfect-information oracle; pairwise differences are examined using Wilcoxon signed-rank tests.
Experimental evaluation setup described in the paper: comparisons to two heuristic baselines and an oracle, and use of Wilcoxon signed-rank tests for pairwise comparisons.
Demand shocks are modeled using two specifications: a mixed profile (half the products follow a uniform demand process and the rest follow a Merton-type jump-diffusion process) and a fully shock-driven profile.
Modeling choices described in the methods: two demand-shock specification setups for simulation experiments.
Policies are learned using Proximal Policy Optimization (PPO) in an actor–critic architecture, with bounded stochastic policies to handle constrained action spaces.
Method description in the paper specifying the use of PPO, actor–critic structure, and bounded stochastic policy parameterization.
The study develops a centralized Hierarchical Reinforcement Learning (HRL) control framework that makes decision timing explicit: replenishment and allocation are optimized weekly, while fulfillment and lateral inventory rebalancing are controlled daily.
Methodological description in the paper: design of an HRL framework with two-level timing (weekly vs daily) for different control decisions.
Algorithmic accuracy alone does not determine value; legitimacy and uptake hinge on people's and process readiness.
Thematic conclusion drawn from interviews, Likert surveys, and document analysis across cases indicating non-technical factors strongly influence uptake despite algorithmic performance metrics. (Sample size not reported.)
The study utilized 3.87 million consumer comments from 127,846 product listings to build and validate models.
Data description reported in paper: 3.87 million consumer comments and 127,846 product listings used.
Molecular representations discussed include string-based methods, topological models, five key categories of Graph Neural Networks (GNNs), 3D-aware Geometric Deep Learning (GDL), emerging Quantum Machine Learning (QML), and Hybrid Quantum-Classical Neural Networks (HQNNs).
Taxonomy and descriptive enumeration of representation classes provided by the review (no empirical comparison or performance claims quantified in the provided text).
The study combines theoretical analysis with quantitative empirical research using survey data from Bosnia and Herzegovina analyzed by regression.
Paper summary states the methodological approach: theoretical analysis plus a quantitative empirical study based on survey data from Bosnia and Herzegovina, analyzed with regression methods. No further methodological details or sample size provided in the summary.
The long-term dynamic effects of AI on resilience remain unverified and require longer-term data.
Authors explicitly state the need for longer time-series data to validate long-term dynamics.
Enterprise-level indicators used in the study do not directly capture supply chain network structure and node dependencies.
Explicit limitation noted by the authors about measurement and scope.
The study's sample is limited to listed manufacturing companies, so conclusions should be applied cautiously to small and medium-sized enterprises (SMEs).
Explicit limitation stated by the authors in the paper.
Mediation and moderation models are leveraged to explore how AI enhances resilience via resource allocation optimization, productivity, and technological innovation, and how conditional factors (e.g., agility) affect these links.
Authors state they used mediation and moderation models on firm-level data to test mechanisms and conditional effects.
The study uses data on A-share listed manufacturing companies from 2011 to 2023 and applies a multi-period difference-in-differences (DID) model to assess AI's impact on SCR.
Methods description provided in the paper summary: sample timeframe and econometric approach explicitly stated.
We introduce a new benchmark QuantSightBench to assess prediction-interval forecasting capability and evaluate frontier models under multiple settings, assessing both empirical coverage and interval sharpness.
Methodological contribution reported in the paper: creation of QuantSightBench and its use to evaluate models on empirical coverage and sharpness (paper describes benchmark and evaluation procedure; specific task/sample counts not given in excerpt).
Technology-driven recruitment encompasses Applicant Tracking Systems (ATS), AI-powered screening, video-based interviews, gamified assessments, and data analytics.
Conceptual description in the paper's introduction/background defining the scope of 'technology-driven recruitment'.
The study employed a mixed-methods research design combining a quantitative survey of 150 HR professionals and recruiters across manufacturing, IT, banking, and education sectors with qualitative case study analysis of four organizations in Chhatrapati Sambhajinagar.
Explicit methodological statement in the paper: quantitative survey (N=150) across specified sectors + qualitative case studies of 4 organizations in Chhatrapati Sambhajinagar.
The study used a mixed-method approach, combining qualitative and quantitative analysis of multiple case studies involving AI applications such as computer vision, robotics, and predictive analytics.
Authors report study design as mixed-method (qualitative + quantitative) applied to multiple case studies examining AI applications (computer vision, robotics, predictive analytics). No numeric sample size reported in the summary.
Future research should prioritize longitudinal and comparative studies to bridge the gap between experimental promise and practical application.
Authors' stated research agenda/recommendation in the review's conclusion.
Findings were synthesized narratively due to methodological heterogeneity.
Methods/results statement in the review explaining narrative synthesis choice because of heterogeneity among included studies.
Risk of bias was assessed using the ROBINS-I tool.
Methods statement in the review specifying ROBINS-I for risk-of-bias assessment.