Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
Existing research has clear gaps: limited evidence from developing-country contexts, insufficient attention to within-occupation heterogeneity, incomplete accounts of psychological mechanisms underlying AI anxiety, and a shortage of rigorous evaluations of reskilling policy effectiveness.
Author's assessment based on the reviewed literature identifying thematic gaps and methodological limitations (critical literature review).
An exploratory evaluation compared unstructured vibe coding, structured prompt engineering, and the Shift-Up approach in the development of a web application.
Paper reports an exploratory evaluation / comparative study described in the abstract; the task context is a web application development exercise comparing three approaches (no sample size reported in abstract).
The First Fundamental Theorem of Welfare Economics assumes that welfare-bearing agents are autonomous and implicitly relies on a binary distinction between autonomy and instrumentality.
Explicit statement in the paper's introduction/abstract describing the theorem's assumptions; conceptual/theoretical textual analysis (no empirical sample).
We evaluate structural validity, semantic alignment, reproducibility, and refinement effort to characterize authoring scalability.
Reported evaluation dimensions in the paper; implies empirical assessments were performed along these axes (details not provided in the abstract).
Hierarchical regression analysis and bootstrapping methods were employed for empirical testing.
Methods section explicitly states use of hierarchical regression and bootstrapping for empirical tests on the survey data.
The study used a three-wave longitudinal survey design collecting matched data from 497 employees.
Methods section states a three-wave longitudinal survey and reports matched data from 497 employees.
We evaluate 20 state-of-the-art LLMs on their ability to predict empirically supported causal directions.
Experimental evaluation: 20 LLMs tested on the benchmark (10,490 triplets, including 1,056 contested instances) to predict empirically verified causal signs.
From 10,490 causal triplets (treatment-outcome pairs with empirically verified effect directions) derived from top-tier economics and finance journals, we identify 1,056 ideology-contested instances.
Construction/extension of the EconCausal benchmark by selecting 10,490 causal triplets from top-tier economics and finance journals and labeling 1,056 as ideology-contested (intervention- vs market-oriented divergence).
The paper contributes by sharpening the concept of management accounting decision quality, distinguishing GenAI from broader digital transformation, and offering a cautious process model grounded in documentary case evidence from leading Chinese manufacturers.
Author-stated contribution in the paper: conceptual refinement and process model based on the three-case documentary analysis.
Because the evidence is drawn primarily from external disclosures rather than direct internal observation, the claims should be read as interpretive analytical inferences rather than as definitive causal proof.
Author's own limitation statement about data sources (external corporate disclosures) and inferential scope.
The study adopts an interpretive multiple-case design and analyzes three major Chinese manufacturing firms - Midea Group, Haier Smart Home, and Dongfang Electric - using official annual and semi-annual reports, corporate disclosures, and recent AI-and-accounting literature.
Explicit methodological statement in the paper: interpretive multiple-case design; data sources listed as official annual and semi-annual reports, corporate disclosures, and literature; sample consists of three named firms.
This is an exploratory and qualitative state-of-practice study grounded in over 30 interviews across four stakeholder groups (large enterprises, small/medium firms, AI developers, and CAD/CAM/CAE vendors).
Methodological statement in the paper describing study design and sample composition.
Key breakthroughs needed include integration with traditional engineering tools and data types, robust verification frameworks, and improved spatial and physical reasoning.
Interviewee-identified requirements compiled from over 30 interviews; stakeholders repeatedly pinpoint integration, verification, and spatial/physical reasoning as priority technical advances.
We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, measuring information aggregation by the log error of the last price.
Statement of experimental design and measurement approach in the paper: laboratory-style controlled experiment, private signals given to agents, log error of last price used to quantify aggregation.
Allowing strategic prompting does not affect information aggregation.
Experimental manipulation that included strategic prompting of AI agents prior to trading; aggregation measured by log error of last price; observed no effect.
Changing the initial price does not affect information aggregation.
Experimental condition varying the initial market price and measuring resulting aggregation performance (log error of last price); reported no effect.
Changing the duration of the market does not affect information aggregation.
Experimental manipulation of market duration in the trading experiment; measured aggregation (log error of last price) across durations and found no effect.
Allowing cheap talk communication does not affect information aggregation.
Experimental condition comparing markets with and without cheap talk communication; aggregation measured by log error of the last price; reported no effect.
As AI reduces the costs of ideation, synthesis, and search, the central bottlenecks of science increasingly shift toward coordination, adjudication, validation, and adaptive steering.
Argumentative/trend claim presented in the paper as motivation for PIM; no empirical time-series or quantitative analysis provided in the paper itself.
The paper formalises crowdsourced R&D and hackathon-type architectures as operational search forms and links these to Causal Problem Modelling (CPM) and the Causal Theoretical Twin Architecture (CTTA).
Conceptual mapping and theoretical linkage between existing crowdsourcing/hackathon models and CPM/CTTA within the PIM framework (theoretical exposition; no empirical mapping or measurement reported).
PIM proceeds through causal problem decomposition, distributed search, real-time evidential updating, contribution traceability, staged validation, and dynamic reprioritisation of candidate solution pathways.
Procedural description of the PIM methodology and its constituent stages in the paper (methodological/theoretical exposition; no experimental implementation reported).
PIM is designed for problem spaces characterised by causal heterogeneity, partial observability, nonlinear interaction, long feedback delays, and distributed expertise.
Methodological design specification within the paper describing the target problem-space features for which PIM is intended (conceptual specification; no empirical testing).
This paper formalises extensions of crowdsourced R&D and hackathon-based research into a general methodology called Probabilistic Innovation Methodology (PIM).
The paper presents a conceptual/theoretical formalisation and names the resulting methodology PIM (no empirical study or sample reported).
The paper includes a companion video demonstrating the approach: https://youtu.be/55Q3lq1fINs.
Statement in paper providing link to companion video.
The physical robot scenario used a 7-DOF robot arm to validate the approach.
Experimental setup description in paper specifying hardware used (7-DOF robot arm).
Prior research typically considers task-level and motion-level adaptation in isolation (task-level methods ignore spatial interference; motion-level methods ignore broader task context).
Literature summary/related work section asserting the separation of prior task-level and motion-level approaches.
RAPIDDS models an individual's spatial behavior (motion paths) and temporal behavior (time required to complete tasks) over multiple cycles.
Description of modeling approach in the paper (method details describing spatial and temporal individual models over multi-cycle interactions).
This paper introduces RAPIDDS, a framework that unifies task-level and motion-level adaptation for human-robot teaming.
Methodological contribution described in paper (framework design and implementation).
This study proposes a framework for evaluating platform ecosystems by their long-term effects on human capital formation and institutional resilience.
Methodological contribution claimed by the paper (development of an evaluative framework); presented as part of the paper's contributions rather than an empirical finding.
A shift-share design finds no detectable effect of early adoption on worker-reported technology-related task restructuring.
Causal-style shift-share analysis using the 2024 EWCS exposure measures to estimate effects of early generative AI adoption on worker-reported changes in technology-related task content; sample >36,600 workers; result reported as no detectable effect.
We compare multiple state-of-the-art agents (e.g., GPT-4o, Llama 3, Qwen2) on metrics assessing tool selection accuracy, faithfulness, and hallucination.
Paper lists evaluated models (GPT-4o, Llama 3, Qwen2) and reports evaluation on metrics including tool selection accuracy, faithfulness, and hallucination across the benchmark.
Our benchmark consists of 100 financial questions.
Paper explicitly states the benchmark contains 100 financial questions.
Under three scenarios (optimistic: 2028-2035; base: 2035-2045; pessimistic: 2045-2060), we specify disconfirmation criteria that would weaken the thesis if observed.
Scenario analysis and specification of disconfirmation criteria by the authors; methodological claim about forecasting structure rather than empirical result.
Converging evidence from history, philosophy, neuroscience, technology, organizational studies, and cultural analysis supports this thesis.
Authors' multidisciplinary literature review and synthesis across the named fields (method: qualitative review); no single empirical dataset or sample size given.
We introduce 'instrumental dissolution' -- loss of institutional-default status while persisting in specialist niches.
Conceptual/theoretical contribution defined by the authors and illustrated via cross-disciplinary examples; no empirical validation sample reported.
Typing's dominance was instrumental, not cognitively necessary.
Argumentative/historical analysis presented in the paper; synthesis of historical and philosophical literature (no empirical sample or experiment reported).
We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews.
Reported evaluation methods and sample in the paper's abstract: log analysis, surveys, and 20 interviews with over 2,200 participants across 116 countries.
Participants were retested individually on the programming tasks after a retention interval of one week.
Statement in abstract describing follow-up retest procedure (one-week retention interval, individual retest).
Participants were incentivized by bonus compensation to balance performance with understanding.
Paper description of participant incentives in methods/abstract; compensation scheme used during experiment.
We conducted a controlled pair programming study with 22 participants who wrote Python code under time pressure in teams of two and individually with GitHub Copilot for 20 minutes each.
Statement of study design in the paper's methods/abstract; controlled pair programming experiment with 22 participants, 20-minute tasks in both conditions (human teammate and Copilot).
In both popular and academic press, concerns are often expressed that AI threatens not only people’s livelihoods but also the meaning they derive from their work.
Observational/literature-commentary claim made in the paper's abstract; references to discourse in popular and academic press (no empirical study or sample reported).
The analysis uses causal discovery methods and integrates scenario-based outcomes, communication analysis, and questionnaire measures.
Paper abstract states that causal discovery analysis was used and that it integrates scenario outcomes, communication analysis, and questionnaire measures.
The study examines user Extraversion and Agreeableness alongside AI design characteristics including Adaptability, Expertise, and chain-of-thought Transparency.
Variables listed in the abstract as the human personality traits and AI design characteristics analyzed.
The study compares two interaction scenario categories: (1) hiring negotiations between human job candidates and AI hiring agents; and (2) human-AI transactions in which AI agents may conceal information to maximize internal goals.
Explicit description of the two scenario categories in the paper abstract; method: experimental / simulation scenarios.
The study includes a parallel human subjects experiment involving 290 human participants.
Statement in paper abstract reporting a human-subjects experiment with 290 participants.
The study uses a purely simulated dataset comprising 2,000 simulations.
Statement in paper abstract describing a simulated dataset of 2,000 simulations; method: simulation experiments.
Algorithmic accuracy alone does not determine value; legitimacy and uptake hinge on people's and process readiness.
Thematic conclusion drawn from interviews, Likert surveys, and document analysis across cases indicating non-technical factors strongly influence uptake despite algorithmic performance metrics. (Sample size not reported.)
The study utilized 3.87 million consumer comments from 127,846 product listings to build and validate models.
Data description reported in paper: 3.87 million consumer comments and 127,846 product listings used.
A randomly sampled coalition of equal size remains largely ineffective at increasing platform spending / wages.
Theoretical comparison in the model between targeted coalitions and randomly sampled coalitions of the same size; analytical results showing limited impact for random coalitions.
We contribute junior–senior accounts on their usage of agentic AI through a three-phase mixed-methods study: ACTA combined with a Delphi process with 5 seniors, an AI-assisted debugging task with 10 juniors, and blind reviews of junior prompt histories by 5 more seniors.
Authors' methodological description of the study design and participant counts as reported in the paper.