Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
The study includes AI adoption audits from 120 organizations.
Methodological statement in the paper specifying the audits sample size.
LLM-generated solutions contain roughly the same number of ideas as participant-generated solutions.
Comparative analysis of idea counts within solutions reported in the paper; phrased as 'roughly the same number of ideas' (no numeric effect size provided in the abstract).
AwareLLM was evaluated in a user study with 20 participants, compared to a standard LLM assistant across multiple tasks.
Experimental methods statement in paper; explicitly reports a user study and sample size.
Using an agent-based simulation of a multi-SKU convenience store environment, the study evaluates deployment efficiency, inventory responsiveness, and managerial cognitive reallocation.
Methodological claim: the paper reports an agent-based simulation experiment in a multi-SKU convenience store context; details such as number of simulations, parameter settings, or statistical results are not provided in the excerpt.
The dominant paradigm for AI agents is an "on-the-fly" loop in which agents synthesize plans and execute actions within seconds or minutes in response to user prompts.
Statement in paper presenting a characterization of current AI agent design; conceptual/observational claim with no empirical data or sample reported.
We thematically analysed twelve semi-structured interviews with SME owners and managers conducted in early 2025 using Atlas.ti, yielding 19 codes grouped into six categories.
Methods statement in the paper describing qualitative sample and analysis procedures.
We examine the interplay between AI adoption, social capital formation, workforce dynamics, and sustainable development in Eastern Macedonia and Thrace (EMT), one of the EU's least developed regions.
Study context and scope as stated in the paper; empirical work conducted in EMT.
Research has concentrated on advanced urban economies, leaving the implications of AI for peripheral small and medium-sized enterprises (SMEs) operating under weak human capital, thin digital infrastructure, and constrained social capital — underexplored.
Statement in the paper contrasting existing research focus (advanced urban economies) with a lack of attention to peripheral SMEs; no empirical sample size for this bibliographic claim reported in the excerpt.
This study conducts an empirical analysis using data on industrial robots from the International Federation of Robotics (IFR) and panel data from 14 sub-sectors of China's manufacturing industry.
Statement in paper describing data and methods: use of IFR robot data combined with panel data covering 14 manufacturing sub-sectors (panel regression framework implied).
The synthesis covers research and practitioner guidance from the years 2023–2025.
Methods statement specifying the temporal scope of sources used for the synthesis.
This paper synthesizes recent research and practitioner guidance (2023–2025) to develop a practical model for designing human–AI collaboration in the financial reporting function (controllership).
Methods section declaration describing scope and approach (literature/practitioner guidance synthesis covering 2023–2025).
We conducted a controlled experiment comparing traditional task-splitting methods with AI-assisted approaches using GitLab Duo.
Methodological statement in the paper reporting a controlled experiment using GitLab Duo; sample size not stated in the provided summary.
The study uses a panel dataset of 35,347 firm-year observations from 2010 to 2023.
Reported sample description in the paper: panel dataset covering 2010–2023 with 35,347 firm-year observations.
AI-assisted decision-making paradigms do not have a significant direct effect on task performance.
Experimental study of 59 pre-service teachers using a two-factor mixed design (between-subjects: AI-assisted decision-making paradigms; within-subjects: human-AI consistency). Data analyzed with Bayesian cumulative link mixed model and structural equation modeling; authors report no significant direct effect.
The authors ran a within-subjects study comparing authoring AD from scratch against editing AI drafts of varying quality.
Explicit methodological statement in the paper (within-subjects study design); sample size not reported in the excerpt.
RefineAD is an editing interface for human revisions (used to compare human editing of AI drafts against authoring from scratch).
Description of the authors' interface/tool used in the study (methodological claim).
GenAD is an AD generation pipeline that incorporates accessibility guidelines and contextual video information.
Description of the authors' system design (methodological claim from the paper).
The study tested Olava Extract against five frontier models.
Method statement in the paper/abstract specifying comparison with five frontier models.
Prevailing metrics, Task Success Rate (TSR) and Agent Handoff F1-Score (HF1), capture only final outcomes or unordered routing decisions.
Conceptual critique presented by the authors; no quantitative validation presented for this claim within the excerpt.
Perceived responsiveness (a functional cue) did not function as a general mediator of anchor type on trust.
Moderated mediation analyses in the randomized experiment (N = 439) found no overall mediation via perceived responsiveness across the full sample.
Methodologically, the work integrates dual eye-tracking, pupillometry, episode-based analysis, and causal inference to capture SSRL as a dynamic, emergent process.
Description of methods and measurement approach across studies: dual eye-tracking, pupillometry (for JME), episode-based analysis, and causal modeling are reported as combined methodology.
The paper reports three eye-tracking studies involving 182 dyads engaged in collaborative debugging tasks.
Stated description of methods in the paper: three eye-tracking studies, total sample of 182 dyads, task = collaborative debugging.
Data analysis utilized regression modeling for performance correlations, time-series analysis for predictive maintenance patterns, and thematic analysis for qualitative interviews.
Paper methods: explicit listing of analytic techniques used (regression, time-series, thematic analysis).
Secondary data encompasses sustainability reports, carbon footprint assessments, and operational performance metrics.
Paper methods: explicit listing of secondary data sources (sustainability reports, carbon footprint assessments, operational metrics).
Blockchain transaction records spanning eighteen months across Nigeria were used as primary data.
Paper methods: explicit statement about 18 months of blockchain transaction records across Nigeria.
The study uses IoT sensor data from forty-five facilities.
Paper methods: explicit statement that IoT sensor data were collected from 45 facilities.
Primary data collection includes structured interviews with supply chain managers.
Paper methods section: primary data described as including structured interviews with supply chain managers (number of interviewees not specified).
The study uses mixed methods involving case studies from twelve multinational companies across the manufacturing, logistics, and retail sectors.
Paper statement of methods: explicit mention of mixed methods and case studies from 12 multinational companies across the three sectors.
The study was a randomized trial of 356 clinicians generating 7,476 trust ratings.
Methods/results reported in paper specifying randomized design, N=356 clinicians, total of 7,476 trust ratings collected.
Prompt-driven generation (even with detailed prompting) fails to address the central problem of architectural complexity management in AI-based software engineering.
Results showing prompting did not prevent code bloat/coupling; conceptual argument reframing the problem toward architecture management rather than prompt engineering.
Neither functional correctness nor detailed prompting mitigates this architectural decay in AI-generated code.
Experimental comparisons reported in the paper where functionally correct outputs and variants produced with more detailed prompting were evaluated for structural quality and showed persistent architectural degradation.
User-defined constraint types maintain usability.
User studies report that despite the additional constraint-typing features, usability remained acceptable (details, metrics, and sample sizes not provided in excerpt).
We conducted a technical evaluation and user studies with general and expert participants.
Paper reports carrying out both a technical evaluation and user studies (methods section). Specific sample sizes not provided in excerpt.
AI learns from both explicit knowledge (papers, documentation, structured databases) and implicit knowledge (reasoning patterns, debugging processes, intermediate steps).
Stated as a conceptual premise in the position paper; no empirical methods, sample, or quantitative data reported.
Perceived usability and satisfaction among participants showed little difference across model sizes.
Reported participant-reported measures (usability and satisfaction) compared across model sizes 3B, 8B, and 70B for N=112 participants; paper states little difference across sizes (no numeric statistics provided in the excerpt).
We examine the performance of humans (N=112) assisted by RAG-assistants compared to LLM-only or LLM+RAG baselines.
Experimental comparison reported in the paper with N=112 human participants across conditions (human+RAG vs LLM-only vs LLM+RAG baseline conditions).
This work evaluates a chatbot-style assistant based on Retrieval-Augmented Generation (RAG) in a realistic multi-turn information-seeking scenario inspired by workplace settings where compliance with local legislation and secure handling of sensitive data are often key.
Reported experimental setup: a chatbot-style RAG assistant evaluated in a realistic multi-turn information-seeking scenario inspired by workplace settings (method description in the paper).
The paper evaluates 'Spec Kit' and 'TDAD' as instantiations of the SGM via a four-month pilot study.
Empirical pilot evaluation reported in the paper; duration specified as four months. Sample size or number of teams/participants in pilot not specified in the summary.
The paper identifies two amplifying mechanisms for PRP: the code review bottleneck and the context window constraint.
Theoretical argumentation in the paper naming two mechanisms that amplify the PRP phenomenon (qualitative explanation).
The paper formally defines PRP with three moderating variables: task abstraction, codebase maturity, and developer experience.
Theoretical/formal definition presented in the paper identifying three moderators; claim is descriptive of the paper's conceptual model.
This paper conducted a multivocal literature review of 67 sources spanning 2022–2026.
Statement of method in the paper describing the literature review (count of sources = 67).
Telemetry across 10,000+ developers shows flat delivery metrics (no improvement in delivery outcomes) despite changes in PR and review behavior.
Observational telemetry across >10,000 developers reported in the paper; described result is no meaningful change in delivery metrics (e.g., delivery throughput, lead time) despite increases in PRs and longer reviews.
A qualitative design was adopted, drawing on 34 semi-structured interviews with project managers across five UK industries.
Qualitative study methods reported in the paper: 34 semi-structured interviews with project managers sampled across five UK industries; Gioia-informed thematic analysis.
A symbolic lifting operator translates simulator trajectories into qualitative descriptors, motion labels, temporal predicates, and structural diagnostics that models interpret across iterative design cycles.
Architectural/methodological contribution described in the paper: a symbolic lifting operator that converts simulator trajectories into higher-level symbolic diagnostics used by LM agents during iterative refinement.
Language model agents explore discrete topologies while numerical optimisers fit continuous parameters.
Methodological description of the system architecture in the paper (division of labor between LM agents for discrete topology search and numerical optimisers for continuous parameter fitting).
There is little empirical exploration of how professionals making high-stakes decisions perceive their agency and level of control when working with genAI systems.
Statement about a gap in the existing literature made by the authors (literature review / framing); no sample size (gap claim).
The literature review employs the PRISMA model to screen, identify, and synthesize available literature on AI, Machine Learning and Deep Learning in promoting managerial productivity and task efficiency.
Methodological statement in the paper's abstract (explicitly states use of PRISMA for screening and synthesis).
Using the Iterated Prisoner's Dilemma (IPD) is an effective scenario to probe cooperative behavior and the influence of visual inputs on VLM decision-making.
Methodological choice described in the paper: experiments were structured around repeated IPD games to operationalize cooperative vs. selfish decisions under visual priming conditions.
This study enriches AI capability research by incorporating engineering perspectives and extends organizational learning theory by examining how AI capability shapes decision-making processes within engineer–AI collaboration contexts.
Authors' theoretical contribution claims in the discussion/conclusion, grounded in their empirical results from the questionnaire study (n=435).
Established scales from authoritative foreign journals were used for measurement, with appropriate translation and verification procedures carried out.
Methods statement that the study used established scales from authoritative foreign journals and performed translation and verification procedures for measures.