Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
The literature on AI-powered emotional intelligence systems is fragmented and insufficiently synthesised.
Authors' assessment based on a systematic literature review conducted following the PRISMA framework (details of databases, search terms, and included studies reported in the paper); exact number of studies not stated in the abstract.
The framework can be operationalized in future empirical research (the article outlines directions for operationalizing the framework).
Methodological/research agenda claim stated in the article's conclusion; describes future empirical operationalization rather than presenting results.
Density-normalized outcomes (e.g., smells per LOC) can mislead when treatment affects system size; raw counts and explicit decomposition are required for causal mining studies of AI tool adoption.
Interpretation and methodological recommendation derived from the observed pattern (unchanged smell counts + increased LOC leading to lower density) in the paper's empirical results.
Per-type estimates and robustness checks (wild cluster bootstrap, Lee bounds, stale-observation sensitivity) corroborate the main pattern; pre-trends are flat (Wald p = 0.90), consistent with the parallel trends assumption.
Placebo and robustness analyses reported in the paper (per-type breakdowns and multiple sensitivity checks) applied to the 151-repository panel; pre-trend test result reported as Wald p = 0.90.
Total architectural smell counts are essentially unchanged after adoption (+1.1%, p = 0.82).
Estimated treatment effect from staggered DiD / Borusyak imputation on total smell counts using the 151-repository panel (74 treated, 77 controls).
The causal effect of adoption on architectural smell density (ASD) was estimated using a staggered difference-in-differences design and the Borusyak imputation estimator.
Methodological claim describing the causal identification and estimation strategy applied to the 151-repository panel.
We mined 151 open-source Java repositories, 74 with detectable agentic AI adoption (identified via configuration files and Co-Authored-By commit trailers) and 77 propensity-matched controls, across a 13-month per-repository window yielding 1,811 monthly Arcan snapshots.
Descriptive dataset and methods statement in paper: 151 repositories (74 treated, 77 matched controls), 13-month windows, producing 1,811 monthly Arcan snapshots.
Causal evidence on the effect of AI coding tool adoption on software architecture is scarce; prior causal work has focused on code-level outcomes (complexity, static analysis warnings) and whether such degradation propagates to architecture-level outcomes remains unknown.
Literature/background statement in abstract asserting gaps in prior work; not an empirical result from this paper's dataset.
We analyze 15,549 agentic PRs from 148 projects in the AIDev dataset.
Descriptive statement of dataset and sample used in the study (paper reports analysis of 15,549 agentic PRs from 148 projects).
From 2024 to 2026, more than 130 articles were submitted to this Special Issue (SI), and only 18 papers were accepted after rigorous peer review.
Editorial report in the paper describing CFP submissions and acceptance counts.
We conduct a qualitative study on a representative sample of 306 non-merged pull requests created or co-authored by the agents mentioned earlier, followed by a quantitative analysis of the reasons for rejection.
Authors' reported methods: qualitative study of a sample of 306 non-merged PRs and subsequent quantitative analysis.
Professional radiologists analyzed chest X-rays with access to state-of-the-art machine learning predictions in this replication setting.
Description of the experimental/contextual setting in the paper: professional radiologists using ML predictions on chest X-rays drawn from Collab-CXR.
The replication uses radiologist assessments from repeated-case designs, which include 68 radiologists and 11,420 paired radiologist–patient–pathology observations.
Direct reporting of study sample and design in the paper (repeated-case design; counts of radiologists and paired observations).
This note leverages the public Collab-CXR data repository described by Moehring et al. (2025) and first analyzed for human-AI collaboration by Agarwal et al. (2023).
Explicit statement in the paper identifying the data source and prior analyses (references provided).
In a production switchback experiment, the offline-trained policy reduces courier-side time costs without degrading customer-facing delivery quality.
Empirical claim supported by production switchback experiment described in the paper; asserts no degradation in customer-facing delivery quality concurrent with courier-side time improvements (no numerical metrics or sample sizes provided in excerpt).
In our setting, the locus of AI bias is not estimation but interpretation.
Overall experiment results: agent coefficient/estimate distributions remained aligned with human consensus and largely unchanged under biased prompts, while final-verdict outcomes were flip-prone under confirmatory prompts (e.g., Claude Code 10%→90%).
Unlike for biased human analysts in the same data, the anti-immigration prior prompt does not shift agents' aggregate estimates or final verdicts.
Comparison of the effect of an anti-immigration prior on human analysts (reported bias) versus agents (20 runs), showing that agent aggregate estimates and final verdict rates remained stable despite changes in methodological decisions.
No agent model exactly matches any human model.
Specification-by-specification comparison showing that none of the agent-generated models (from 20 executions) are identical to any human analyst's model in the many-analysts baseline.
Both agents' effect estimates remain broadly aligned with the human consensus.
Comparison of effect estimate distributions from Claude Code and Codex (20 runs each) to the human many-analysts consensus; reported alignment/broad agreement between agent estimates and human consensus.
At the design layer, Codex matches human methodological diversity.
Comparison of methodological specifications produced by Codex (20 independent executions) to the many-analysts human baseline; reported similarity in diversity metrics between Codex outputs and human analysts.
We run 20 independent executions of Claude Code and Codex on a prominent immigration and social-policy problem and compare them against a many-analysts human baseline.
Experimental method described in the paper: 20 independent runs/executions of each agent model (Claude Code and Codex), compared to an existing many-analysts human baseline.
After screening, 35 studies were included in the thematic synthesis and supplemented by official regulatory and industry documents.
Review screening result reported in the paper: number of included studies = 35; supplementation by regulatory and industry documents stated.
A structured search protocol was designed for Scopus, Web of Science, PubMed, IEEE Xplore, and Google Scholar covering January 2016 to May 2026, English-language records only.
Methods statement in the review describing the databases, date range, and language restriction used for the systematic search.
The implementation literature on AI for pharmacy inventory and pharmaceutical supply chains remains dispersed across pharmacy operations, operations research, health informatics, and supply chain analytics.
The review's thematic synthesis of the searched literature (review methods described below) identified studies across these disciplinary areas.
Specification, reference implementation, conformance suite, and worked examples are available at: https://github.com/BrightbeamAI/chap
Claim of artifact availability hosted on GitHub (URL provided) as part of the paper's resources.
Two protocol standards address adjacent concerns: MCP standardises agent access to tools and data, and A2A standardises agent-to-agent interoperability.
Factual claim referencing existing standards (MCP and A2A) and their scopes; no citations or supporting documentation included in the provided excerpt.
Production deployments are no longer one human supervising one model; they are multi-human, multi-agent collaborations that cross teams, time zones, and trust boundaries.
Stated as a general characterization of modern production deployments; no quantitative data or case counts provided in the excerpt.
Retrieval augmentation and scientist persona prompting yield only marginal gains.
Ablation/augmentation experiments comparing baseline LLM outputs to versions augmented with retrieval or scientist-persona prompting, showing only small improvements in judged quality.
6,749 scientists returned 25,139 sets of ratings on novelty, empirical feasibility, probability of being true, and favorability of adoption.
Reported study participation and rating counts: 6,749 respondents providing 25,139 rating sets on specified dimensions.
We invited authors of 121,640 recent preprints across biology, medicine, chemistry, and the social sciences to judge follow-up ideas that large language models (LLMs) generated from the context and puzzles of their own papers.
Study recruitment described in paper: invitations sent to authors of 121,640 recent preprints across multiple fields (biology, medicine, chemistry, social sciences).
The findings provide empirical insights for managing employee wellbeing and refining human resource strategies during organizational digital transformation.
Authors' stated implications in the discussion, based on the reported empirical associations and moderation results from the survey of 411 employees.
The study draws on the Conservation of Resources Theory and the Cognitive Appraisal Theory of Stress to explain how AI application influences employees' job insecurity via resource gain and resource threat mechanisms.
Theoretical framing stated in the introduction and discussion explaining the mechanisms (resource gain vs. resource threat) underlying the observed U-shaped association.
Data were collected via mixed online and offline questionnaires: 453 questionnaires were distributed (242 online, 211 offline); 449 were returned (242 online, 207 offline); following validity screening, 411 valid questionnaires were retained (219 online, 192 offline), yielding an effective response rate of 90.73%.
Reported survey administration and response counts provided in the methods section of the paper.
Devil's Advocate (DA) is an AI assistant that critiques the human's initial ideas, whereas Dialectical Inquiry (DI) provides alternatives and synthesizes a resolution.
Conceptual/definitional claim in the paper describing the operationalization of DA and DI for the experiments.
This research empirically compares DA and DI in AI contexts.
Paper reports experimental comparison between AI behaviors implementing Devil's Advocate (DA) and Dialectical Inquiry (DI) across the studies.
Both studies examine benefit (information elaboration) and cost (cognitive load) pathways when AI supports SDM.
Paper explicitly frames both studies to measure information elaboration as a benefit pathway and cognitive load as a cost pathway; stated measurement plan in methods.
Study 2 tests mind-shaping interventions through user strategy training.
Study design described in the paper: a second experiment (Study 2) manipulating user strategy training (mind-shaping) to evaluate effects on SDM processes and outcomes.
Study 1 tests tool-shaping interventions by comparing three AI bot prototype conditions (Information-only, DA, DI) against a control treatment.
Study design described in the paper: randomized/controlled experiment (Study 1) with four conditions (three AI prototype conditions plus control).
The 'do no harm' property is confirmed empirically.
Abstract states empirical confirmation in simulations and applications; specifics (e.g., datasets, sample sizes) not included in abstract.
Including AI predictions as covariates has a 'do no harm' property: the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative.
Stated theoretical property in the paper and described as empirically confirmed in simulations and applications (per abstract).
Raw blind-panel decision quality is similar for A and B (7.01 vs. 6.96).
Blind-panel scoring of generated reports from agents A and B; panel size and panel methodology not specified in abstract.
The value of an in-band cooperative deny signal (Recuse Signal) is an empirical question: it was previously unmeasured and the paper measures whether compliant LLM agents honor such a signal.
Motivation and framing in the paper; they position their controlled experiment as the measurement addressing this previously unmeasured question.
We searched seven databases (plus backward and forward citation searching) and synthesised 13 empirical studies published between 2018 and 2025.
Methods reported in abstract: PRISMA-ScR scoping review with a preregistered protocol; explicit count of included studies and publication date range.
Self-evaluated creative performance remained unchanged when using GenAI.
Same experiment with 82 participants; authors report no significant difference in self-evaluated creative performance between GenAI users and controls.
Each of the four published papers used in the experiments contained an error that I helped identify or correct.
Author statement that the 4 papers each contained an error; author involvement in identification/correction is asserted.
I conducted experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGPT) to check the correctness of four published papers in economic theory.
Author reports running direct experiments: prompted listed models to check 4 published economic-theory papers.
From Codeforces histories we build an AI-prompt signature characterised by more first-attempt acceptances and fewer attempts and retries, consistent with AI-assisted practice.
Empirical construction from CF submission histories (pattern: increased first-try accepts, fewer retries). Method: analysis of historical submission logs; sample size not stated in abstract.
The International Collegiate Programming Contest (ICPC) and the International Olympiad in Informatics (IOI) prohibit AI under proctoring and admit entrants through qualification rounds, whereas online Codeforces (CF) contests are unproctored and open to all.
Descriptive factual claim about contest rules and formats (institutional description in paper); based on contest rules and organizational formats referenced by authors.
We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels.
Paper's stated evaluation methodology: operator feedback + production question set, graded by humans and automated panels.
Over 100 participants collaborated with one of four frontier models (Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7) on a long-horizon coding task lasting around five hours.
Study description: experimental participants (reported as "Over 100 participants") each paired with one of four named models on a ~5-hour coding task designed to mimic real-world workflows.