Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The study collected data from 207 entrepreneurial businesses (including SMEs, startups, and knowledge-based businesses) using a structured questionnaire and analyzed the data using Partial Least Squares Structural Equation Modeling (PLS-SEM) with SmartPLS 3.
Structured questionnaire administered to a sample of 207 entrepreneurial businesses; analysis conducted with PLS-SEM (SmartPLS 3) as reported in the paper.
The study analyzes the influence of artificial intelligence, financial technology, economic performance, monetary policy, financial development, and governance quality on the growth of G7 countries over 2000–2024 using the Method of Moments Quantile Regression (MMQR).
Statement in paper specifying use of Method of Moments Quantile Regression on G7 countries during 2000–2024. Implied panel sample: 7 countries × 25 years ≈ 175 country-year observations (if annual, balanced panel).
We conducted preregistered experiments in two tasks (a sentiment-analysis task and a geography-guessing task) to study whether user characteristics influence the effectiveness of AI explanations.
Preregistered experimental studies described in the paper; two distinct tasks (sentiment-analysis and geography-guessing). (Sample sizes and additional procedural details are not provided in the excerpt.)
The paper empirically analyzes the algorithm-automated versus human decision-making debate using the AST and STS theoretical lenses.
Theoretical analysis and empirical synthesis across the reviewed studies (n=85), explicitly stated use of AST and STS frameworks to interpret findings.
To address the duality of benefits and harms, the paper proposes a dynamic Human-in-the-Loop (HITL) model that reconciles algorithmic determinism with normative HRM demands.
Conceptual/theoretical contribution presented in the paper (proposed HITL model based on synthesis of findings and theory).
There is substantial heterogeneity in effects (I^2 = 74%), indicating variability across studies.
Meta-analytic heterogeneity statistic reported in the paper (I^2 = 74%).
This study analyzes 28 papers (secondary studies and research agendas) published since 2023.
Systematic literature review conducted by the authors of secondary studies and research agendas; sample size explicitly reported as 28 papers; timeframe specified as 'since 2023'.
Three contributions are presented: the Agentic AI Framework (AAF 3.0); a cross-domain synthesis formalising the inverse evidence–complexity relationship; and a phased sociotechnical roadmap integrating governance sequencing, reimbursement reform, and equity safeguards.
Descriptive claim about the paper's outputs. These contributions are stated in the abstract as the study's deliverables based on the narrative review and synthesis of 81 sources.
Agentic AI is defined as autonomous, goal-directed systems capable of multi-step workflow coordination.
Definition provided by the authors within the paper (conceptual framing used for the review).
This structured narrative review of 81 sources (2020–2025) evaluates whether Agentic AI ... can support structural adaptation in ageing health systems.
Methodological statement in the paper: the study is a structured narrative review of 81 sources from 2020–2025.
The framework is depicted across organization areas with primary focus on strategic management and workforce decision-making and secondary focus on finance, operations, and marketing.
Descriptive claim based on the conceptual framework and its mapping to organizational domains within the paper. No empirical application or case studies reported.
This paper outlines a Human–AI Collaborative Decision Analytics Framework integrating five overlapping layers: data, AI analytics, business analytics interpretation, human judgment, and feedback learning.
Presentation of a conceptual framework developed by the authors (conceptual/modeling contribution). No empirical validation reported.
The results presented in the paper are based on a literature recherche, an analysis of individual tasks across different occupations (conducted within Erasmus+ projects), and discussions with trainers/educators.
Methodological statement from the paper; indicates the types of evidence used. The abstract does not provide numbers for analyzed tasks, the number of occupations, details of Erasmus+ projects, or counts of trainers/educators consulted.
Neither time constraints nor LLM use significantly change strategic foresight in the startup evaluation task.
Null findings reported from the same experimental comparisons in the 2 × 2 design (N = 348): no statistically significant effects of time constraints or LLM use on the strategic foresight outcome.
The study employed a 2 × 2 experimental design manipulating time constraints and LLM use.
Explicitly reported experimental design in the paper: two factors (time constraints, LLM use) crossed to form four conditions in the startup evaluation task.
The study used a sample of N = 348 participants.
Reported sample size in the paper's experimental study (startup evaluation task); participants across the 2 × 2 experimental design totaled 348.
The paper identifies key research gaps and proposes a future research agenda focused on human–AI interaction, organizational governance, and ethical accountability.
Conclusions/recommendations from the conceptual meta-analysis (paper-generated research agenda; no empirical testing reported in abstract).
This study presents a conceptual meta-analysis of interdisciplinary literature on AI-augmented decision-making in organizations.
Methodological statement of the paper (the paper itself is a conceptual meta-analysis); no primary empirical sample reported in the abstract.
Research has insufficiently modeled joint distributional outcomes and environmental performance, and lacks integrated evaluation of AI-enabled sustainable finance under heterogeneous disclosure regimes.
Review-level identification of methodological gaps across the surveyed literature (authors' synthesis of existing studies and their limitations).
There is a shortage of long-horizon causal evidence on non-linear coupling between digitalization and decarbonization, limiting robust policy inference.
Meta-level assessment in the review noting gaps in existing empirical literature (review authors' synthesis of the field; claim about research availability rather than primary data).
Competency mapping involves identifying and aligning the critical skills, knowledge, and abilities required for specific job roles.
Definition provided in the paper (conceptual).
A stratified random sampling method was employed to select a representative sample of 500 IT employees, based on a pilot study constituting 0.50 percent of the total population.
Sampling description provided in the methods section: stratified random sampling, sample size = 500, pilot study size referenced as 0.50% of population.
The study analyzes data from the period 2021 to 2023 using Multiple Regression Analysis as the principal analytical technique.
Methods statement provided in the paper (timeframe and analytical method).
The primary objective of this research is to examine the impact of AI adoption on competency mapping practices in the IT sector.
Explicitly stated research objective in the paper.
The study employs the Difference-in-Differences (DiD) method to estimate AI impacts on online labor markets over time.
Methodological statement in the abstract specifying the use of Difference-in-Differences for empirical identification; implementation details (controls, parallel trends checks, sample size) are not given in the abstract.
The Act instituted a rigid seven-percent per-country cap that allocates the same number of visas to India (population of 1.4 billion) as to Iceland (population of 400,000).
Statutory per-country cap (7% rule in the INA) combined with publicly available country population figures for India and Iceland; claim about identical allocation follows directly from the 7% rule.
The Immigration Act of 1990 established a ceiling of 140,000 employment-based green cards annually.
Statutory fact derived from the Immigration Act of 1990 and the Immigration and Nationality Act (INA) provisions setting employment-based annual numerical limits.
Python code and data required to replicate the results are provided in the paper's appendix.
Author statement that 'Python code and data for replication are included in the appendix.'
The empirical analysis uses a smooth-transition local projection model applied to U.S. productivity and EPU data.
Methodological statement in the paper describing the estimation approach and the data inputs; replication materials (Python code and data) are included in the appendix.
This study uses panel data from 30 Chinese provinces (2011–2022) and estimates a spatial simultaneous equations model using the Generalized Spatial Three-Stage Least Squares (GS3SLS) approach.
Described methodology in the paper: panel dataset covering 30 provinces over 2011–2022 (12 years), spatial simultaneous equations estimated by GS3SLS.
Deterministic automated verifiers provide objective pass/fail checks for task success.
Methods section: verifiers are deterministic and automated, enabling objective evaluation of whether an agent's trajectory accomplished the task.
Scale of experiments: seven agent–model configurations and 7,308 execution trajectories were used to compute pass rates and deltas.
Reported experimental scale in Methods: 7 agent–model configurations and a total of 7,308 agent execution traces collected and analyzed across tasks/conditions.
Each task was evaluated under three conditions: (1) no Skills, (2) curated (human-authored) Skills, and (3) self-authored (model-generated) Skills.
Experimental protocol described in Methods: three-arm evaluation per task across the SkillsBench benchmark.
SkillsBench benchmark: evaluates 86 tasks spanning 11 domains with deterministic, automated verifiers.
Dataset and benchmark description in the paper: SkillsBench contains 86 tasks across 11 domains and uses deterministic pass/fail verifiers for objective evaluation.
Research should prioritize dynamic, task-based models that include transitional frictions, heterogeneous agents, and sectoral structure to better measure AI exposure and impacts.
Methodological recommendation grounded in the paper's theoretical critique of static occupation-level automation metrics and noted empirical gaps.
Timing uncertainty and measurement challenges make forecasting the pace and scale of AI-induced employment change inherently uncertain.
Methodological limitations section noting uncertainty in AI adoption speed and difficulties mapping capabilities to tasks and predicting new occupation emergence.
Research agenda: there is a need for causal studies on AI’s impact on accounting labor demand and firm performance, analyses of distributional effects across firm sizes and industries, and evaluation of regulatory frameworks for reliable, interpretable AI in financial reporting.
Author-stated research priorities drawn from gaps identified in the literature review; not an empirical finding.
Policy implications include workforce retraining, standards for AI auditability and transparency, and regulation balancing innovation and controls (privacy, fraud prevention).
Policy recommendations based on identified risks and barriers discussed in the paper rather than empirical policy evaluation.
For stronger causal evidence, recommended empirical methods include difference-in-differences on adopting firms vs. controls, matched samples, and randomized pilots for particular tools, supplemented by qualitative interviews.
Methodological recommendations stated in the paper (not an empirical finding); no implementation/sample reported in the abstract.
Actionable research priorities include running larger-scale field trials linking game use to observed land-use and economic outcomes, developing validation protocols for game-backed models against empirical on-farm data, studying heterogeneity of impacts, and designing incentive mechanisms that leverage game-demonstrated profitability co-benefits.
Synthesis-driven recommendations based on identified evidence gaps—specifically the predominance of small-scale/qualitative studies and lack of long-term/causal evidence.
Rigorous economic evaluation (RCTs, quasi-experiments) is needed to quantify how game-enhanced DSTs affect investment, land-use choices, emissions outcomes, and farm incomes.
Chapter recommendation grounded in observed gaps: the literature lacks sufficiently rigorous causal impact evaluations; current evidence is largely qualitative or observational.
The empirical strategy uses baseline panel regressions with standard controls (e.g., firm size, performance, leverage) and fixed effects to estimate the AI → pay relationship.
Methods section describing regression specifications including firm controls and fixed effects applied to the A-share firm panel.
Data consist of a panel of Chinese A-share listed companies covering 2007–2023.
Data description in the paper specifying the sample period and population (A-share listed firms, 2007–2023).
The firm-level AI application indicator is constructed via textual analysis of corporate disclosures (e.g., filings/annual reports) to capture AI application intensity.
Methodological description in the paper describing text-based construction of an AI application indicator from corporate disclosures for listed firms in the 2007–2023 sample.
Empirical validation of the integrated Kondratieff–Schumpeter–Mandel framework requires firm-level adoption and profitability data, sectoral investment series, and cross-country comparisons using panel methods and identification strategies (e.g., diff-in-diff, IV).
Methods/limitations section recommendation (explicitly states no single micro-econometric identification strategy was reported and outlines required data/methods).
The three frameworks (Kondratieff, Schumpeter, Mandel) are complementary: Kondratieff frames periodicity, Schumpeter provides micro-mechanisms of innovation-driven change, and Mandel foregrounds socio-political constraints and distributional outcomes.
Conceptual integration and comparative theoretical analysis (qualitative synthesis).
Kondratieff's framework is useful for identifying broad periodicities (recurring phases of expansion and stagnation) in capitalist development but is less specific about microeconomic mechanisms.
Theoretical review of Kondratieff literature and conceptual assessment (qualitative).
No new laboratory measurements or datasets are reported in the paper; the approach is methodological and conceptual rather than empirical.
Methods section and explicit statements within the paper noting absence of new data; verifiable by reading the paper.
These operators are presented as conceptual/theoretical bridges rather than immediately quantifiable laboratory units.
Explicit methodological statement in the paper emphasizing interpretive/theoretical intent; no empirical operationalization reported.
Policy recommendations include: invest in open metadata standards; fund pilot programs to evaluate ROI (earnings, placement, employer satisfaction); require model governance and periodic external audits for AI-assisted curriculum tools; and support smaller providers via shared infrastructure or accreditation hubs.
Explicit policy recommendations in paper (prescriptive).