Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
We evaluate four mechanisms to enable cooperation: (1) repeating the game for many rounds, (2) reputation systems, (3) third-party mediators to delegate decision making to, and (4) contract agreements for outcome-conditional payments between players.
Description of experimental design / mechanisms evaluated in the study across four social dilemmas; details on implementation and sample sizes not provided in the excerpt.
The paper provides lessons for scaling regression automation and enabling effective human-AI teaming in Agile settings.
Stated contribution of the paper (synthesis of lessons from the industrial case study).
The Copilot was integrated with Hacon's CI pipelines and operates asynchronously as a 'silent AI teammate', producing candidate scripts for human review.
System integration and deployment description within the case study (implementation detail reported in the paper).
We conducted an exploratory industrial case study of the Hacon Test Automation Copilot, an agentic AI system that generates system-level regression test scripts from validated specifications using retrieval-augmented generation and a multi-agent workflow.
Methodological claim: description of the study design and the system; the paper reports a single industrial case study at Hacon (a Siemens company).
SAFI measures LLM performance on text-based representations of skills, not full occupational execution.
Methodological caveat stated by the authors clarifying the scope and limits of SAFI.
We propose an AI Impact Matrix that positions skills into four quadrants: High Displacement Risk, Upskilling Required, AI-Augmented, and Lower Displacement Risk.
Conceptual/interpretive framework introduced by the authors; described in text as proposed by the paper.
Legitimate accountability is axiomatized through four minimal properties: Attributability (responsibility requires causal contribution), Foreseeability Bound (responsibility cannot exceed predictive capacity), Non-Vacuity (at least one agent bears non-trivial responsibility), and Completeness (all responsibility must be fully allocated).
Paper presents an explicit axiomatization listing these four properties as definitions/axioms forming the normative criteria for legitimate accountability.
Collective behaviour is characterised through interaction graphs and joint action spaces.
Paper specifies interaction graphs and joint action spaces as part of the formal model (definitions and formal structure).
Autonomy is characterised through a four-dimensional information-theoretic profile (epistemic, executive, evaluative, social).
Paper defines autonomy as a 4-dimensional information-theoretic profile (conceptual/mathematical definition within the formal model).
Using a strictly algorithmic baseline (mathematical bottleneck aggregation), we calculate Relative Occupational Automation Indices (OAI) for the U.S. labor market based on the DWA-level scores.
Method and calculation claim: algorithmic baseline aggregation applied across the 923 occupations / 2,087 DWAs to produce OAIs mapped to the U.S. labor market. Specific aggregation formula referenced but not numerically detailed in the excerpt.
We deconstructed 923 occupations into 2,087 Detailed Work Activities (DWAs).
Explicit data processing claim in the paper: mapping of 923 occupations to 2,087 DWAs for analysis.
A life insurance system integrated into an industry partner mobile app was tested in two experiments.
Paper reports two experiments running the ARQuest-enabled life insurance system inside a partner mobile app; experimental setup is stated though sample sizes are not provided in the excerpt.
We evaluate the architecture through a controlled experiment (600 runs across five industries: FinTech, Insurance, Healthcare, Vietnamese Banking, and Vietnamese Insurance).
Controlled experiment reported in the paper: 600 runs across five named industries (experimental setup reported in abstract).
The framework is calibrated with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, and implemented in computer vision.
Calibration and empirical implementation using O*NET, a domain expert survey (n=3,778), and GPT-4o task decompositions; applied to computer vision tasks.
We introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level.
New metric proposed in the paper (entropy-based task complexity) and mapping procedure from accuracy to substitution ratio; implemented in the framework.
Costinot and Werning (2023) develop a sufficient-statistic approach and find optimal technology taxes of 1–3.7% on robots.
Citation reported in the paper summarizing Costinot and Werning (2023)'s quantitative sufficient-statistic estimate.
Guerreiro et al. (2022) characterize optimal Mirrleesian tax system with automation and find that robot taxes should be transitional—high when incumbent workers cannot retrain, converging to zero as new cohorts adjust skill investments.
Citation reported in the paper summarizing Guerreiro et al. (2022)'s theoretical result on transitional robot taxes.
If labor becomes economically redundant, the policy focus shifts from steering innovation to redesigning public finance and redistribution (e.g., new tax instruments, redistribution mechanisms).
Theoretical scenario analysis in the paper with references to related works (Korinek and Juelfs 2024; Korinek and Lockwood 2026).
Evaluation is carried out under three frozen context configurations (diff only: config_A; diff with file content: config_B; full context: config_C) enabling systematic ablation of context provision strategies.
Methodological description: three fixed context configurations defined and used for ablation experiments.
We performed an extensive evaluation of 37 state-of-the-art Vision-Language Models on MultihopSpatial.
Empirical evaluation described in the paper listing the number of models evaluated (37).
Economic evaluations of GLAI should account for end-to-end risk externalities (error propagation, institutional trust, rights impacts), not only short-term productivity gains.
Methodological recommendation grounded in conceptual synthesis of technical, behavioral, and legal risks; normative argument rather than empirical result.
Generative Legal AI (GLAI) systems are built on token-prediction (LLM) architectures rather than formal legal-reasoning architectures.
Conceptual and technical analysis in the paper distinguishing GLAI from other legal-tech; literature synthesis on common LLM architectures. No original empirical dataset or sample size—qualitative/technical review.
Through a thematic review of existing research, the authors identified recurring themes about incentive schemes: their components, how researchers manipulate them, and their impact on research outcomes.
Authors' stated method and findings: thematic review (the scope/number of reviewed papers not specified in excerpt).
A critical aspect of conducting human–AI decision-making studies is the role of participants, often recruited through crowdsourcing platforms.
Claim based on the authors' thematic literature review noting participant sourcing practices (specific studies and counts not given in excerpt).
Researchers conduct empirical studies investigating how humans use AI assistance for decision-making and how this collaboration impacts results.
Statement summarizing the research landscape; supported implicitly by the authors' thematic review of existing empirical studies (number of studies not specified in excerpt).
Returns to AI are heterogeneous across firms; estimating treatment effects requires attention to selection, complementarities, and dynamic adoption pipelines.
Methodological argument referencing treatment-effect literature and observed firm heterogeneity; supported by conceptual examples rather than a single empirical treatment-effect estimate.
Human-only and AI-assisted teams performed similarly on most outcomes.
Comparison across outcome measures from the randomized experiment; summary statement indicates parity on most measured tasks except for detection of major coding errors.
We randomly assigned 288 researchers to 103 teams working under three conditions (human-only, AI-assisted, AI-led).
Experimental design reported in paper: randomized assignment of 288 researchers into 103 teams across three experimental conditions.
An undirected configuration change in a prior iteration produced zero impact, illustrating the cost of iterating without diagnosis.
Reported observation from an earlier iteration in the case study where a non-diagnostic change had no measurable effect.
The paper combines findings from information systems research, organizational behavior studies, and artificial intelligence literature through its analysis of recent empirical and theoretical studies conducted between 2021 and 2026.
Methodological description provided in the paper (literature synthesis covering 2021–2026).
Results are robust to state-by-year and industry-by-year fixed effects.
Robustness checks reported in paper that include state-by-year and industry-by-year fixed effects with results stated to hold.
Where AI can perform tasks independently, we find no significant employment effect.
Heterogeneous DiD estimates showing null (statistically non-significant) employment coefficients for occupations/industries where AI can perform tasks independently.
We examine aggregate effects using administrative data covering essentially all U.S. employers in a difference-in-differences design exploiting occupational AI exposure across industries and states.
Statement in paper describing data and empirical strategy: administrative data covering essentially all U.S. employers; difference-in-differences design exploiting occupational AI exposure variation across industries and states.
Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner.
Result from the in-person pilot (N = 62) comparing originality scores between participants partnered with GPT-4 versus human partners under matched time limits; reported as statistical equivalence in the paper.
We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage.
Experimental setup description specifying four task-lifecycle stages and use of one generated skill for each stage.
A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content.
Additional control experiment (token-matched content) reported in supplement, with run count and comparison results described.
The total spread across variants is only 1.2 percentage points.
Reported range/variation in performance metrics across all skill variants in the ablation experiment.
Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396.
Statistical hypothesis tests comparing Full and ablated-skill variants to task-only prompting; reported minimum p-value threshold.
The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs.
Description of experimental design and aggregate counts reported in the paper (tasks × model configurations × providers → total run count).
We find no reliable improvement from full generated skills over No-Skill prompting.
Empirical comparison between Full generated-skill prompting and No-Skill (task-only) prompting across the study's evaluation tasks and model configurations; statistical testing reported.
We survey 1,250 arXiv papers (2024-2026).
Systematic literature survey conducted by the paper; explicit statement of sample size and date range in the abstract.
From that coded sample the authors built a causal model of 26 constructs and 67 relationships (64 directed, 3 contested).
Reported model construction from the coded sample as stated in the abstract.
The authors filtered that corpus and coded a stratified random sample of 3,100 documents with an LLM-assisted pipeline.
Reported sampling and coding procedure stated in the abstract.
We collected 38,709 grey-literature documents (engineering blogs and Reddit threads) and filtered to those substantively about code review.
Reported data-collection procedure and corpus size stated in the abstract.
The study uses difference-in-differences linear regressions on 2023 and 2024 KBO season data to identify the causal impact of ABS adoption by player status.
Methods statement in paper: difference-in-differences linear regressions; sample sizes reported as n = 148 batters and n = 112 pitchers.
High-status pitchers' performance remains unaffected by ABS adoption.
Difference-in-differences linear regressions using KBO 2023 and 2024 season data for pitchers (n = 112); paper reports no detectable change for high-status pitchers.
The Korea Baseball Organization (KBO) officially implemented the Automated Ball-Strike System (ABS) in 2024.
Paper statement of policy change and use of 2023 and 2024 KBO season data; presented as factual background to the natural experiment.
The periodization of US macroeconomic productivity cycles was refined by identifying the new stages 'pandemic and adaptation phase' and 'artificial intelligence phase'.
Calculation of AAPC indices for 1947–2025 and retrospective comparative analysis leading to refinement of periodization and naming of new stages.
Eight distinct macroeconomic cycles of productivity change in the United States from 1947 to 2025 are identified.
Secondary data analysis of aggregated US Bureau of Labor Statistics series for 1947–2025; long-term average annual rates of productivity change (AAPC) computed using the index method and geometric mean growth rate; comparative analysis to identify cycle breaks.
Survey data were collected from firms located in major Chinese cities (Beijing, Shenzhen, Xi’an, and Zhengzhou), resulting in 750 valid responses for analysis.
Reported survey sampling and data collection in the paper; explicit statement of cities sampled and number of valid responses (750).