Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
A time-lagged survey design was adopted to reduce common method bias.
Authors explicitly state they used a time-lagged survey design intended to reduce common method bias in data collection.
Data on AI capability, algorithmic transparency, decision-making quality, and innovation performance were collected from 435 participants using the Credamo data platform.
Methods section statement that a questionnaire survey was conducted via Credamo and that data for the listed constructs were collected from 435 participants.
All artifacts associated with this study are publicly available at https://zenodo.org/records/18489222.
Statement in the paper providing a Zenodo link to artifacts.
This review identifies key research gaps and provides recommendations for future research and practice.
Authors' discussion and conclusion sections synthesizing gaps and offering recommendations based on the mapping results.
Satisfaction, Performance, and Efficiency are the most frequently investigated SPACE dimensions, whereas Communication and Activity remain underexplored.
Frequency counts and synthesis across the 39 included studies mapped to SPACE dimensions as reported by the authors.
Only 15% of the reviewed studies extend beyond three SPACE dimensions.
Authors' coding of included studies against the SPACE framework with reported proportion.
90% of the reviewed studies adopt a multi-dimensional perspective by examining at least two SPACE dimensions.
Authors' coding of included studies against the SPACE framework, yielding the reported proportion.
This paper is a systematic review and mapping of 39 peer-reviewed studies published between January 2014 and December 2024 that examine the impact of LLM-assistants on software developer productivity.
Authors conducted a systematic review and mapping exercise covering peer-reviewed studies within the stated date range; the paper reports the count of included studies as 39.
Long-running agents accumulated thousands of sequential decisions; continuously active agents reached 6,000+ prompt-state-action cycles.
Agent activity traces showing sequential decision counts per agent (trace-level telemetry).
The system consumed roughly 70B inference tokens across the deployment.
API/inference telemetry reporting total token usage.
More than 5,000 ETH was deployed by agents during the experiment.
Accounting of ETH held/deployed by agent-controlled vaults during deployment.
Agents executed about $20M in trading volume over the deployment.
Aggregate trading-volume accounting from the bounded onchain market during deployment.
The deployment produced roughly 300K onchain actions.
Onchain transaction logs aggregated over the deployment.
The system produced 7.5M agent invocations during the deployment.
System invocation logs reporting total agent calls across the deployment.
DX Terminal Pro was deployed for 21 days with 3,505 user-funded agents trading real ETH in a bounded onchain market.
Deployment logs and system telemetry from a 21-day field deployment reporting the number of user-funded agents.
We evaluate SecMate in a controlled study with 144 participants and 711 conversations.
Reported experimental study sample and conversation counts in the paper.
Given the limited sample size, the results should be interpreted as exploratory.
Authors explicitly note limited sample size (20 decks) and label findings exploratory.
Reliability (stability across repeated runs) varies substantially across models, with ICC values ranging from 0.240 to 0.930.
Paper reports interclass correlation coefficient (ICC) analysis of model output reliability across runs, giving a range of ICC values from 0.240 to 0.930.
To account for stochastic variation in outputs, each model pair was evaluated five times under identical conditions.
Paper states that each model pair was run five times under identical conditions to distinguish one-off variation from persistent tendencies.
Each model evaluated 20 real startup pitch decks spanning multiple industries and funding stages.
Paper reports a controlled simulation design in which each model assessed 20 real pitch decks (sample of 20 decks).
The study used three leading models—GPT-4o, Claude 3.5 Sonnet, and DeepSeek-V2.
Explicit statement in the paper describing the experimental subjects: three named LLMs were evaluated.
We run over 1,100 games with over 16,000 private conversations totaling 15.2 million tokens and over 150,000 player actions.
Dataset and experimental log statistics reported in the paper.
We run AI-only games and conduct a user study pitting human players against AI opponents.
Method statement in the paper describing experiments with both AI-only and human-vs-AI games.
Players have asymmetric objectives and negotiations are non-binding, allowing alliances to form and break as players' short-term interests align and diverge.
Specification of game mechanics and rules in the paper (design features of C2C).
We introduce Cooperate to Compete (C2C), a multi-agent environment where players can engage in private negotiations while competing to be the first to achieve their secret objective.
Description of a newly developed environment (paper introduces the game and its rules/design).
Semantic search maintained comparable inter-rater agreement while reducing chart abstraction time.
Clinical utility evaluation reports that inter-rater agreement was comparable between semantic-search-assisted abstraction and clinician-performed chart review.
The authors optimized embedding model and chunking strategy using a physician-authored benchmark dataset.
Methods: experiment described as optimization of embedding model and chunking using a physician-authored benchmark dataset.
The system uses instruction-tuned qwen3-embedding-0.6B embeddings, stores vectors in a managed database with storage-optimized indexing, maintains full-text metadata in a low-latency key-value store, and operates within a HIPAA-compliant governance framework.
Methods description of system architecture and governance provided in the paper.
We deployed a semantic search system indexing 166 million clinical notes (484 million vectors) from 1.68 million patients.
Paper reports a production deployment at a large children's hospital and gives exact index counts: 166 million clinical notes, 484 million vectors, 1.68 million patients.
We formalize the distinction between compensatory and non-compensatory decision regimes and define a pre-execution legitimacy boundary.
Theoretical formalization presented in the paper (definitions and conceptual framework). No empirical evidence or sample size provided.
Most existing approaches implicitly assume that once a decision is produced, it is eligible for execution.
Author assertion / conceptual critique of existing approaches presented in the paper (no empirical test reported).
Most existing approaches to AI safety, risk management, and governance focus on post-hoc validation, probabilistic risk estimation, or certification of model behavior.
Author statement summarizing the literature / prior work in AI safety and governance (conceptual claim in the paper's introduction). No empirical survey or sample size reported.
We develop a formal model in which institutions choose the scale of automation, the degree of codification, and safeguards on iterative use.
Methodological statement: the paper presents a formal/theoretical model specifying institutional choice variables (model description rather than empirical result).
The welfare consequences of genAI can be organized by a two-dimensional taxonomy: the strength of the incentive to perform the task without AI, and the severity of model collapse.
Analytical organization derived from the theoretical model presented in the paper (conceptual taxonomy based on model parameters; no empirical sample reported in abstract).
We develop a parsimonious model of behavior in collaborative interactions in which individuals can either exert human effort, rely on genAI, or refrain from work altogether.
Methodological claim: authors present a formal theoretical model with the specified choice set (model description in paper; no empirical sample reported in abstract).
Learning-based control offers a more adaptive alternative, but it remains unclear whether such methods... can sustain hours of reliable operation, deliver consistent quality, and behave safely around people on a live production line.
Framing of a research gap in the paper's introduction; no primary experimental data presented here (statement of uncertainty motivating the study).
Die Studie basiert auf einer wiederholten Querschnittsbefragung lizenzierter Beschäftigter einer außeruniversitären Forschungseinrichtung.
Autorenangabe im Abstract: wiederholte Querschnittsbefragung (survey) unter lizenzieren Beschäftigten der untersuchten Forschungseinrichtung; methodische Beschreibung im Abstract.
We use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows.
Methodological contribution described in the paper: a unified amortized computational framework applied to eight Shapley variants, evaluated under latency constraints typical of operational workflows.
No formulation improved objective analyst performance.
Controlled/empirical experiment reported in the paper evaluating eight Shapley variants with professional analysts in the fraud-detection environment; performance measured over 3,735 case reviews.
Standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility.
Empirical comparison in the paper between quantitative metrics (sparsity, faithfulness) and human-judged clarity/decision-utility across the datasets and analyst reviews; based on the authors' large-scale evaluation.
We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews.
Experimental methods reported in the paper: evaluation across four risk datasets and a fraud-detection environment with professional analysts; stated sample of 3,735 case reviews.
A central issue is how humans interpret the algorithm's choice of features, which affects the design and evaluation of highlighting policies.
Framing and motivation in the paper: conceptual claim motivating the formal models and analysis (theoretical/argumentative).
We illustrate our framework in a calibrated empirical exercise based on the American Housing Survey.
An empirical/calibrated exercise using data from the American Housing Survey reported in the paper; the claim is that the framework is illustrated empirically (data-based demonstration).
Humans may interpret the algorithm's choice of features in different ways: a sophisticated agent correctly conditions on the selection rule, while a naive agent updates only on revealed feature values and treats the selection event as exogenous.
Conceptual/behavioral modeling in the paper that defines two agent-types (sophisticated vs naive) and analyzes their distinct inference processes (theoretical/modeling).
Highlighting can be modeled as a constrained information policy that selects a small number of features to reveal.
Modeling framework developed in the paper: formal definition of highlighting as an information policy with a feature-selection constraint (theoretical/modeling).
We study this question using 10,659 matched human-agent pairs from Moltbook, a social media platform where each autonomous agent is publicly linked to its owner's Twitter/X account.
Descriptive statement of the study dataset reported in the paper: dataset of 10,659 matched human-agent pairs from Moltbook with public linkage to owner's Twitter/X account.
The paper proposes a conceptual framework linking AI adoption to employability and role transformation, mediated by skill adaptation, continuous learning, and organizational readiness.
Author-proposed conceptual framework presented in the review paper (theoretical linkage based on literature synthesis).
This study takes food delivery riders as the research object and analyzes the dilemma of labor relations determination under AIGC.
Methodological statement in the paper specifying the chosen subject of analysis (food delivery riders); this is an explicit description of the paper's scope rather than an empirical finding.
The paper develops an interdisciplinary conceptual framework that integrates insights from economics, management theory, and digital governance to characterize algorithmic enterprises.
Methodological claim about the paper's approach; stated in abstract as the paper's contribution (conceptual framework built from interdisciplinary literature).
Future research should strengthen cross-national comparisons, longitudinal tracking, and interdisciplinary collaboration to support development of a technology governance framework that balances efficiency with equity.
Author recommendation based on identified research gaps in the literature review (prescriptive/recommendation).