Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The review followed PRISMA guidelines.
Methods statement in the paper indicating PRISMA adherence.
After screening, 10 studies met the inclusion criteria.
PRISMA-style screening result reported in the review (records screened and included).
A comprehensive search across Scopus, Web of Science, IEEE Xplore, and ScienceDirect yielded 260 records.
Systematic search following PRISMA guidelines reported in the paper; databases searched explicitly listed.
The LLM fallacy is situated within existing literature on automation bias, cognitive offloading, and human–AI collaboration, but is distinguished as a form of attributional distortion specific to AI-mediated workflows.
Conceptual positioning and literature synthesis in the paper; claim is analytic rather than empirically tested in the abstract.
Less attention has been given to how LLM usage reshapes users' perceptions of their own capabilities.
Literature gap claim from the paper's review of prior research on model reliability, hallucination, and trust calibration; no quantitative synthesis or meta-analysis reported.
The system evaluation was performed in a deployed multi-tenant enterprise application across three conditions: manual operation, unconstrained AI with safety layers disabled, and full bounded autonomy.
Method description in the paper's evaluation section: deployment context (multi-tenant enterprise app), three experimental conditions, and 25 scenario trials spanning seven failure families.
The review focuses on the 2020–2025 period for studies of AI application in financial auditing.
Stated scope/timeframe of literature included in the review.
Article selection was conducted using the Scopus (Q1–Q4) and Sinta (1–2) databases based on predefined inclusion and exclusion criteria, resulting in a final sample of 15 articles.
Stated data sources and selection procedure in the Methods section; final sample size explicitly reported as 15.
This study employs a Systematic Literature Review (SLR) method following the PRISMA 2020 protocol.
Stated methodology in the paper: explicit use of SLR and PRISMA 2020 protocol.
We evaluate AIBuildAI on MLE-Bench, a benchmark of realistic Kaggle-style AI development tasks spanning visual, textual, time-series and tabular modalities.
Evaluation methodology described in paper (benchmark selection and task modalities).
We surveyed 860 Microsoft developers to understand where they want AI support, and where they want it to stay out.
Primary empirical method reported in the paper (survey) with sample size explicitly stated as 860 Microsoft developers.
Developers spend roughly one-tenth of their workday writing code.
Statement reported in the paper (abstract). No sample-size or measurement method for this specific statistic provided in the abstract.
We examine 12 tasks across two practical settings: an AI consultancy providing solutions to business problems and an AI software team developing software products.
Description of experimental design and sample reported in the paper (method section): 12 tasks, two practical settings.
We tested 9 frontier models on BTB.
Abstract states that nine frontier models were evaluated using the benchmark.
Completing a BTB task takes bankers up to 21 hours, underscoring the economic stakes of successfully delegating this work to AI.
Reported time-to-complete statistic in abstract (claimed maximum of 21 hours per task); implies measurement of human task completion time by bankers.
BTB requires agents to execute senior banker requests by navigating data rooms, using industry tools (market data platform, SEC filings database), and generating multi-file deliverables including Excel financial models, PowerPoint pitch decks, and PDF/Word reports.
Benchmark design specifications reported in abstract describing the tasks and artifact types agents must produce.
BankerToolBench (BTB) is an open-source benchmark of end-to-end analytical workflows routinely performed by junior investment bankers.
Paper describes BTB design and explicitly states it is open-source and targets end-to-end workflows for junior investment bankers.
We collaborated with 502 investment bankers from leading firms to develop an ecologically valid benchmark grounded in representative work environments.
Reported collaboration/sample size stated in abstract: 502 investment bankers involved in benchmark development.
The full model, including all 11 analytical tabs, is made publicly available to facilitate replication and independent sensitivity testing.
Paper states that the full model and all 11 analytical tabs are publicly available.
A sensitivity analysis shows that the high-skill capture rate and the pace of friction decay are the two parameters with the greatest influence on the aggregate result.
Paper reports results of a sensitivity analysis identifying parameter importance; explicitly names high-skill capture rate and friction decay pace as most influential.
AI coverage scores are sourced from Massenkoff and McCrory (2026) and mapped to NAICS industries using employment-weighted averages derived from BLS Occupational Employment and Wage Statistics data for 2023.
Citation to Massenkoff and McCrory (2026) for theoretical LLM task coverage across SOC groups and explicit statement that mapping used employment-weighted averages from BLS OES 2023.
The core formula multiplies six inputs: base GDP, labor share, AI coverage, productivity gain percentage, adjusted adoption rate, and a skill-weighted capture rate.
Model specification in the paper describing the multiplicative core formula and listing the six inputs.
The two case firms demonstrated contrasting approaches to implementing AI in recruitment.
Findings and case descriptions comparing the two firms' AI recruitment strategies and levels of implementation (n = 2 firms; interviews with 22 participants).
The research contributes by shifting focus to under-researched non-Western workplace settings, particularly technologically advancing Middle Eastern economies like Qatar.
Paper's stated contribution and scope: focus on Qatari organisations and Middle Eastern context.
Four key themes emerged from the data: (1) process optimisation through AI integration, (2) subjectivity in AI-powered recruitment, (3) recruitment strategies in the age of AI, and (4) strategic investments in AI.
Findings: thematic analysis identified these four themes from interview data (n = 22) across the two case firms.
Thematic analysis was used to identify patterns and relationships within the interview data.
Methods: analysis section reporting use of thematic analysis framework.
Data were collected through semi-structured interviews with twenty-two participants across various organisational roles and hierarchical levels.
Methods: semi-structured interviews reported with total participants n = 22 across roles/levels.
The research investigated two prominent Qatari firms with contrasting AI recruitment implementation approaches.
Methods / case selection: two firms were selected and contrasted on their AI recruitment approaches (number of firms = 2).
The study employed an interpretivist philosophy and a case study design.
Methods section: explicitly states interpretivist philosophy and case study design.
All participants had access to the same AI tool; the experiment varied only the structure surrounding its use (behavioral vs cognitive scaffolding vs unstructured).
Experimental design description in the paper: common AI tool provided to all participants; randomization/assignment varied only the scaffolding around AI use.
These results are observational and reflect a single-operator dataset without controlled comparison.
Author statement in the paper describing study limitations.
There is a significant research gap in comparative understanding of generative AI's impact across developed and developing economies; differences in infrastructure, labour markets, and skill distributions may lead to uneven outcomes.
Review observation that the included literature lacks sufficient comparative studies across country-development contexts (explicitly noted as a gap in the paper).
This systematic literature review synthesised findings from 40 empirical and conceptual studies published between 2020 and 2025 using the PRISMA framework (search across Google Scholar and Dimensions.ai), yielding 3,252 database records plus 8 hand-searched studies, of which 40 met the inclusion criteria.
PRISMA-style structured literature search reported in the paper: database search (Google Scholar, Dimensions.ai) returning 3,252 records, 8 hand-searched records, 40 studies meeting inclusion.
We instantiate this vision in a controlled study (n=36) comparing the gaze-aware AI assistant to a text-only LLM assistant.
The paper reports running a controlled user study with sample size n=36 directly comparing the gaze-aware assistant against a text-only LLM assistant.
Exploratory innovation does not show a significant direct association with long-term competitive performance.
PLS-SEM results from the survey of 104 Portuguese B2B managers reporting a non-significant direct path from exploratory innovation to performance.
Data were analyzed using partial least squares structural equation modeling (PLS-SEM) implemented in SmartPLS 4.
Methods section statement in paper indicating use of PLS-SEM and SmartPLS 4 for data analysis.
The empirical analysis is based on a questionnaire survey administered to 324 respondents from Romanian organizations operating in IT, services, industry, and public administration.
Questionnaire survey described in paper; sample size explicitly stated as 324 respondents from Romanian organizations across IT, services, industry, and public administration.
On document intelligence (DocILE), our Code Factory variant matches Direct LLM on key field extraction (KILE: 80.0%).
Empirical evaluation reported on DocILE dataset of 5,680 invoices; KILE metric reported at 80.0%.
We evaluate on two task types: function-calling (BFCL, n=400) and document intelligence (DocILE, n=5,680 invoices).
Statement in paper specifying dataset/task types and sample sizes used in evaluation.
We validate this principle through a controlled experiment on log format token economy across four conditions (human-readable, structured, compressed, and tool-assisted compressed).
Controlled experiment described in the paper comparing four log-format conditions (human-readable, structured, compressed, tool-assisted compressed); exact sample size not reported in the abstract.
For six decades, software engineering principles have been optimized for a single consumer: the human developer.
Historical/position claim asserted in the paper (conceptual/literature-based argument), no empirical sample or quantitative test reported.
A series of robustness checks were conducted to ensure the reliability of the conclusions.
Paper statement that multiple robustness checks were performed in support of the main DiD findings (e.g., alternative specifications, placebo tests, etc. implied).
The study uses China's National New-Generation Artificial Intelligence Innovation Development Pilot Zone (NAIDPZ) as a quasi-natural experiment and applies a staggered difference-in-differences (DiD) model on panel data of 267 Chinese prefecture-level cities from 2007 to 2023.
Paper statement of research design: staggered DiD model applied to panel data covering 267 prefecture-level Chinese cities over 2007–2023, treating NAIDPZ as quasi-natural experiment.
Through a causal decomposition that automates one side of agent communication, we separate cooperation failures from competence failures, tracing their origins through agent reasoning analysis.
Method described in the paper: causal decomposition approach that automates one side of communication and analyzes agent reasoning to attribute failures (methodological claim; abstract mentions the approach but gives no sample size or quantitative metrics there).
Capability does not predict cooperation.
Comparative experimental results reported in the paper showing different models with different capability levels achieving substantially different collective cooperation outcomes (specifically comparing OpenAI o3 and o3-mini).
We build a multi-agent setup designed to study cooperative behavior in a frictionless environment, removing all strategic complexity from cooperation.
Methodological description in the paper: design and implementation of a multi-agent experimental setup intended to remove strategic complexity (no sample size or quantitative detail reported in the abstract).
The empirical basis of the study is industry data from the Bureau of National Statistics of the Republic of Kazakhstan for 2020–2024.
Statement in the paper specifying the data source and years used for calibration of the model.
The study's methodological framework integrates the Bass model of innovation diffusion, an expanded production function with endogenous technological progress and the task-oriented Acemoglu–Restrepo approach, plus a multi-criteria system of industry prioritisation.
Description of the paper's modelling approach in the methods section; model components identified explicitly in the paper.
We evaluated EcoAssist through benchmarks of 500 websites and a controlled study with 20 developers.
Explicit methodological statement in paper: benchmark sample size = 500 websites; user study sample size = 20 developers.
Functional correctness (test-based correctness) exhibits negligible statistical association with design satisfaction.
Statistical analysis reported in experiments comparing test pass (functional correctness) and design-satisfaction labels produced by verifier; paper states negligible association.