Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
We study shared-workspace human-AI teams using the Collaborative Gym environment with DiscoveryBench tasks.
Methodological description in the paper stating the experimental environment and task suite used.
We ran 1,482 sessions in our experiments.
Statement in the paper reporting total experimental sessions (Collaborative Gym + DiscoveryBench).
The article presents a comprehensive framework spanning macro-level governance principles and micro-level interaction typologies, illustrated through case examples from telecommunications, retail, insurance, and consumer products sectors.
Descriptive/methodological claim in the paper indicating the presence of a framework and sectoral case examples (no quantitative sample sizes reported for cases).
The article draws on Deloitte's 2026 Global Human Capital Trends survey of over 3,000 business leaders across 15 countries.
Methodological/data source statement explicitly provided in the paper.
The robot is initialized with the selected episodic memory before a new collaboration episode begins (i.e., memory reuse occurs at episode start).
Design and procedure reported in paper: authors state that the robot is initialized with the selected memory prior to new episodes as part of the intervention tested in the experiment.
Historical collaboration patterns (CPs) can be represented as knowledge-graph episodic memories and used for reuse via graph representation learning with a node-classification objective to identify a representative and effective memory.
Methodological description in the paper: authors construct knowledge-graph episodic memories from prior CPs and apply graph representation learning with a node-classification objective to select memories for reuse; no external validation beyond the study's experimental tests.
The dataset contains large high-resolution images of dimensions 1280x959 and 960x703, which increase the complexity of the annotation task.
Dataset description in the paper specifying image dimensions (1280x959 and 960x703) and noting annotation complexity due to image size.
LLM guidance did not increase the total number of victims saved (no increase in total victims saved relative to baseline).
Same experimental comparison (two LLM-guided conditions vs no-LLM) in the simulated SAR environment; behavioral measure of total victims saved reported.
A 2015-2017 backward extension (224 firms, 601 observations) supplies pre-treatment data and provides evidence against pre-existing upward-trend confounds in SG&A-to-revenue.
Additional panel extension covering 2015-2017 with 224 firms and 601 firm-year observations, used to test pre-trends.
South Korea exemplifies national-scale under-augmentation: high human capital (H), substantial AI (A), but low convergence capacity (C) produce phi = 0.
Case/example presented in the paper as an illustrative national example (descriptive/case-study evidence).
This systematic literature review (SLR) synthesizes empirical studies concerning AI and the implications of these changes on labor skills across all sectors between 2017 and 2025.
Methodological claim in the paper describing its scope and timeframe (SLR of empirical studies, 2017–2025). No numeric count of included studies provided in the excerpt.
Across the subsequent ~150 sessions after deploying Baseline-Log Physical Separation, no recurrence of Index Sickness was observed.
Authors' observational report from Bang-v3 following deployment: subsequent ~150 collaborative sessions with no observed recurrence.
We used action research methods in a real software project (Bang-v3) spanning approximately one month and 391 collaborative sessions.
Paper statement of study design: action research in Bang-v3, duration ~1 month, 391 collaborative sessions.
During preparation of the dissertation, I used generative AI tool ChatGPT for limited language assistance (grammar correction, stylistic refinement, and improving clarity); the intellectual content is entirely the author's own.
Author statement in the dissertation (declaration of use of ChatGPT for language assistance).
The paper proposes an integration framework covering use case suitability, autonomy levels, technical integration, governance, security, employee enablement, and measurable impact.
Paper presents a proposed framework (descriptive; the existence of the proposal is internal to the paper).
Unlike traditional automation or conversational AI, agentic systems can interpret goals, plan multi-step tasks, access tools, interact with enterprise systems, and execute workflows with varying degrees of autonomy.
Descriptive definition and capability listing in the paper (conceptual/technical description). No empirical validation provided.
Agentic AI marks a new phase of enterprise automation.
Author's high-level claim in the paper (conceptual/position statement). No empirical data or sample reported.
A qualitative analysis of 384 stratified patches informs a syntactic taxonomy of eight oracle signal categories.
Manual qualitative coding/analysis of a stratified sample of 384 test-file patches, resulting in an eight-category syntactic taxonomy.
We conduct an empirical study of 86,156 test-file patches from 33,596 agent-authored PRs across 2,807 GitHub repositories produced by five coding agents: OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code.
Dataset collection and descriptive statistics reported by the study (counts of patches, PRs, repositories, and agents analyzed).
Rather than providing empirical validation, this paper outlines the scope of LLM consumer behavior and identifies open research questions related to alignment, preference representation, and market dynamics.
Explicit methodological statement in the paper describing scope: conceptual/theoretical outline and an agenda of open questions rather than empirical results.
The study is grounded theoretically in signaling theory and legitimacy theory to interpret consumer responses to AI-use disclosure.
Theoretical framing stated in the paper referencing signaling theory and legitimacy theory as foundations for hypotheses about how disclosure affects consumer/funder behavior.
This study analyzes Kickstarter projects using an LLM-assisted text classification approach combined with entropy balancing.
Methods statement in the paper specifying use of large language model–assisted classification to detect AI-use disclosures in project texts and entropy balancing as a covariate adjustment technique.
The study recruited participants (N = 123) in a controlled environment to write career plan essays for paired biographical profiles differing only in gender under three conditions: no AI assistance, neutral LLM assistance, or gender-biased LLM assistance.
Experimental recruitment and design as reported in the paper (explicitly states N = 123 and three experimental conditions).
A neutral prompt does not induce gender-differentiated language in LLM-generated essays.
Comparison of LLM outputs under a neutral prompt showing absence of gender-differentiated language (method described in paper). Sample size for these model outputs not specified in the summary.
The production technology is multiplicative and cognitive capital functions as collateral that determines the return to AI adoption.
Analytic model specification and derivation inside the paper (formal mechanism). No empirical data.
The model features two state variables per agent, cognitive capital and cognitive debt.
Formal theoretical model presented in the paper (model specification). No empirical sample; analytic construction.
A pilot deployment in Newham's secure environment evaluated operational performance relative to manual workflows.
Paper reports a pilot deployment and an operational evaluation comparing DOMUS to existing manual workflows (method: pilot deployment; specifics such as duration or sample size not stated in provided text).
Prose features that statistically distinguish AI from human text do not predict which human text gets accused as AI.
Matched-control test comparing accused and matched non-accused comments; analysis showed features known to separate AI vs human text lack predictive power for accusation status among human-authored texts.
A placebo vocabulary of pre-2022 inauthenticity terms (shill, astroturf) did not rise in the same way.
Comparative vocabulary/time-series check using historical (pre-2022) inauthenticity terms as a placebo; these terms did not show the same growth pattern in the dataset.
We ran a matched-control test comparing accused versus non-accused parent comments.
Method description: a matched-control experimental/observational test comparing comments that were accused of AI use to matched comments that were not accused.
We performed speech-act coding of 300 confirmed accusations of AI use.
Methodological statement: a manual / coded speech-act analysis conducted on a set of 300 accusations confirmed as accusations of AI use.
We used LLM judgment on 7,500 sampled accusations of AI use.
Methodological description: a sample of 7,500 accusations was judged/classified using one or more large language models.
We analyzed 25 million comments from Hacker News and Reddit (2023-2026).
Descriptive statement of dataset and time window used for analysis; platform-level scrape/collection of comments totaling 25 million across Hacker News and Reddit (2023-2026).
All experiments reported were preregistered.
The paper's abstract states the experiments were preregistered (series of four preregistered experiments and separate preregistration for the persuasion tournament).
The human comparators in the experiments included laypeople, winners of a separately preregistered four-round online persuasion tournament, professional canvassers, and world championship debaters.
Description in the paper's abstract listing the types of human persuaders used as comparators across experiments.
After coaching, expert humans could tie an AI that was constrained to respond at human speeds and with human-length messages.
Reported follow-up experimental condition in which the AI was artificially constrained (speed and message length) and, after coaching, expert humans matched its persuasion performance.
Across the experiments there were n = 18,978 conversations from 6,923 people.
Aggregate sample sizes reported in the paper's abstract (explicitly states the numbers).
The authors evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels on LabOSBench.
Paper reports an evaluation suite covering multiple model classes and evaluation granularities (subtask and end-to-end); specific model identities and counts are not provided in the excerpt.
LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection.
Explicit specification in the paper of benchmark breadth: 96 subtasks and 8 simulators including the listed workflow stages.
Scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment.
Conceptual claim in the paper describing the nature of scientific-instrument operation; not supported by empirical data in the excerpt.
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems.
Statement in paper framing the problem; based on literature/field observation rather than reported experiments or data in this excerpt.
The study analyzed three-wave survey data from 312 employees from China.
Three-wave longitudinal survey; sample described as 312 employees in China (N=312) used for empirical analyses.
A public tabular operational sanity check tests the associated budget-allocation implication.
Empirical/operational claim: the authors ran a public tabular sanity check to evaluate budget-allocation implications (details and sample size not provided in the abstract).
We test the mechanism in controlled synthetic environments, large-scale A-share factor discovery, and symbolic-regression benchmarks.
Empirical claim about the experimental scope: controlled synthetic experiments, empirical large-scale A-share factor discovery analysis, and symbolic-regression benchmark experiments (datasets and sizes not specified in the provided text).
This study presents a critical synthesis of 50 peer-reviewed articles (2019–2025).
Stated methodological description in the abstract: a critical synthesis of 50 peer-reviewed articles covering 2019–2025.
Those preliminary experiments do not establish behavior preservation, scaling economics, or verified-change cost.
Authors' explicit limitation statement following the preliminary QLoRA experiments.
The review focuses on three core dimensions of impact: employee attitudes (job satisfaction, motivation, adaptability), workplace behaviours (performance, creativity, technology adoption), and organisational dynamics (leadership, trust, team cohesion).
Stated scope and focus areas in the abstract describing the review's analytical framework.
The study contributes a structured framework that clarifies the role of emotional AI in organisational contexts and outlines actionable, scalable strategies for real-world application.
Authors claim to have developed and presented a framework and strategies as part of the review paper (a descriptive/conceptual contribution rather than empirical evidence).
The study identifies key patterns, methodological trends, and underexplored areas in research on emotional AI systems in organisational contexts.
Authors report findings of their comparative analysis of the state-of-the-art literature (exact patterns/trends and counts are detailed in the full review).
This study follows the PRISMA framework to conduct a systematic evaluation and comparative analysis of the state-of-the-art literature on emotional AI in organisations.
Statement of methods in the abstract that the review used PRISMA; implies structured search, screening, and selection procedures described in the full paper.