Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
We conduct the first large-scale study of human oversight in AI coding sabotage.
Authors state they ran a large-scale user study; described as the first such study focused on human oversight in AI coding sabotage (methodological claim).
Traditional software and agentic systems are distinct: in traditional software code is the carrier of decision logic, whereas in agentic systems code is ephemeral tooling used by an LLM-driven reasoning loop.
Formalization and conceptual definitions developed in the paper (first-principles formal distinction; no empirical sample size reported).
For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
Historical/descriptive claim presented in the paper's framing and literature review; citation of longstanding software engineering practices (qualitative, no empirical sample size reported).
We implement a two-stage processing architecture separating document-level extraction (Stage 1) from claim-level synthesis (Stage 2).
Implementation description in paper: architecture design and pipeline stages described by the authors.
The research is grounded in the Resource-Based View (RBV) and Dynamic Capabilities Theory (DCT) to explain how technological and managerial resources contribute to organizational performance.
Author statement in the paper describing the theoretical framework (RBV and DCT) used to frame the study.
The study adopts a quantitative research design and analyzes collected data using Partial Least Squares Structural Equation Modeling (PLS-SEM).
Author statement in the paper describing research design and analytical method.
Digital Leadership did not demonstrate a statistically significant direct effect on Employee Productivity (β = -0.094, p = 0.275).
Reported quantitative result from the study using PLS-SEM; β and p-value provided in the paper showing a non-significant direct effect. Sample size not reported in the excerpt.
We scored over 2.1 million twin responses on 500 participants and 183 held-out questions.
Reported evaluation counts in the paper: 2.1M responses, 500 participants, 183 held-out questions.
The construction-method grid covers three open-weight LLMs, five cumulative information depths ranked by normalized Shannon entropy, two embedding methods, and two reasoning modes.
Paper's experimental design specification (methods section).
We construct detailed individual-level twins from the German Socio-Economic Panel (SOEP) and evaluate them across a 3 × 5 × 2 × 2 construction-method grid.
Methodological description of the study: experimental construction and evaluation on SOEP data.
A large-scale empirical study on Harvey LAB used 12,510 agent trajectories.
Paper states an empirical study run on Harvey LAB with a sample described as 12,510 agent trajectories.
AI-assisted feedback does not reduce time per character (i.e., it does not increase time cost per unit of feedback).
Time-per-character was measured in the randomized field experiment; authors report no reduction (no increase in time per character) associated with the AI-assisted drafts. Student-level/completion-level data from the experiment (n=88); 11 TAs.
AI-assisted feedback does not negatively affect student usefulness ratings.
Measured student ratings of usefulness in the randomized field experiment; authors report no negative effect of the treatment on these ratings (no significant decrease reported). Student-level sample n=88; 11 TAs.
Two-stage field experiments in healthcare prescription messaging encompassed 693,139 patient visits in total.
Paper statement of total sample size across Stage 1 and Stage 2.
Stage 2 (Tool-Augmented Agentic AI) autonomously extracted principles from Stage 1 data and generated 17 new message variants tested on 248,448 patient visits.
Study design and reported results from Stage 2 of the two-stage field experiment described in the paper.
Stage 1 (Human + Chatbot) produced 13 message variants and was tested on 444,691 patient visits.
Study design details reported in the paper describing the two-stage field experiment.
The empirical analysis is based on panel data of new energy vehicle firms in the Yangtze River Delta from 2001 to 2023.
Dataset description provided in the paper's abstract/introduction indicating the time span and regional coverage.
R&D expenditure does not constitute a significant mediating channel between artificial intelligence and firms' new quality productive forces.
Mediation analysis using the panel data and constructed indicators; reported nonsignificant mediation effect of R&D expenditure (no sample size or statistics reported in excerpt).
The system was evaluated on OMH-Polyglot, a multilingual coding benchmark spanning Turkish, Arabic, Chinese, and code-switched specifications.
Experimental evaluation reported in the paper using the OMH-Polyglot benchmark.
Brain privacy has both personal and social attributes; its protection therefore implicates individual interests and technological development.
Normative/legal argumentation and conceptual analysis presented in the paper (no empirical data reported).
The experiment was run twice: a first run with unrealistically loud injections, and a second run with signals rescaled to a physically motivated SNR range.
Protocol described in paper explicitly states two runs with different injection SNR scalings (one 'unrealistically loud', one physically motivated).
Both agents received identical written specifications and identical compute resources.
Methodological statement in paper specifying that both agents were given the same written spec and the same shared computing infrastructure.
The pipeline comprised power spectral density estimation from raw Einstein Telescope simulated noise, geometric template bank generation, matched filter recovery of 100 binary black hole signal injections, automated results generation, and large language model-assisted production of a manuscript formatted in the style of Physical Review D.
Protocol description in paper; matched filter recovery included 100 injected signals (explicitly stated).
We compared two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a simple end-to-end gravitational wave data analysis pipeline on a shared computing infrastructure without human intervention.
Experimental design described in paper: two named agents were given identical written specifications and identical compute resources and executed the full pipeline autonomously.
The audit samples 2,000 runs over a design space of 10 personas x 8 prompts x 3 model configurations x N=10 reps, with the two OpenAI cells at full 8-prompt coverage and the Anthropic sonnet-4.6 / low cell at 4-prompt coverage.
Stated audit design and sample counts in paper (method section describing factorial design and coverage of model/prompt cells).
The paper evaluates the proposed architecture using the outcome metric 'time-to-insight'.
Methodological statement in the paper listing evaluation metrics.
The paper evaluates the proposed architecture using the outcome metric 'time-to-find'.
Methodological statement in the paper listing evaluation metrics.
The paper evaluates the proposed architecture using the outcome metric 'data product adoption'.
Methodological statement in the paper listing evaluation metrics.
In the first acquisition the acquirer pursued a disruptive 'rip-and-replace' strategy for the target’s proprietary ERP system.
Empirical observation from the paper's comparative case study of two consecutive acquisitions of the same digital target (qualitative case evidence).
This paper contributes a large-scale empirical dataset involving 57,954 essays from 10,195 students across 120 schools over two years.
The paper explicitly states the dataset size and coverage in the abstract: 57,954 essays, 10,195 students, 120 schools, two-year period.
We distill our findings into a meta-design and four design principles (DPs), grounded in kernel theories, for systems where human contextual intelligence and algorithmic recognition must coexist.
Design contribution presented in the paper (meta-design artifact and four DPs derived from the study).
We developed a collaborative forecasting system that leverages semantic processing using large language models (LLMs) to solve the 'cold-start' problem for novel menu items while preserving human agency via override mechanisms.
Description of system design and implementation produced during the ADR project (practice-driven abductive approach).
This paper reports on a 9-month action design research (ADR) project at a German financial services firm.
Explicit methodological description in the paper (study duration and organizational context).
We examined how different degrees of embodiment affect team performance and conversational dynamics in a real-life escape room; teams were composed of either three humans or two humans and an artificial agent (a Box, an Avatar, or a hyper-realistic humanoid).
Experimental field study reported in the paper: a real-life escape room experiment comparing team compositions (3 humans vs. 2 humans + agent of three embodiment types). Sample size not reported in the provided text.
To the best of the authors' knowledge, no prior study has examined the psychological mechanism through which algorithmic management shapes employee voice and silence behaviour outside of gig economy and platform work contexts.
Author claim based on literature review (stated gap in existing research).
Estimation accuracy depended only weakly on message volume, indicating that more text alone does not guarantee better inference.
Analysis reported in the paper examining the relationship between message volume and estimation accuracy; described as a weak dependency.
We employ the Gemini API to generate reward function logic and weights across three refinement rounds rather than performing per-step inference.
Methodological description in abstract: use of Gemini API to generate reward logic and weights; three rounds of refinement.
We deploy a Soft Actor-Critic (SAC) agent in CityLearn v2 for experiments.
Methodological description in abstract: SAC agent used within CityLearn v2 environment.
We use four empirically grounded occupant profiles from the ASHRAE Global Thermal Comfort Database II (13,440 votes).
Dataset citation and sample size reported in abstract: ASHRAE Global Thermal Comfort Database II with 13,440 votes; four occupant profiles derived from it.
We ran 24 matches pairing 23 expert humans with 16 AI agents, capturing 387 delegation and 1440 adoption decisions.
Author-reported experimental setup and counts from the study (24 matches; 23 human experts; 16 AI agents; counts of delegation and adoption decisions).
The model introduces the 'Sciencepreneur' as the central human archetype in agentic R&D.
Conceptual/design claim within the HARMONY artifact presented in the paper.
Evidence also includes pattern matching with documented agentic R&D deployments.
Methodological statement in the paper claiming pattern matching with documented agentic R&D deployments (unspecified number/source).
The study includes a foresight scenario analysis projecting four plausible 2040 R&D futures to stress-test design choices.
Methodological statement in the paper describing a four-scenario foresight analysis.
Empirical evidence for the design is triangulated from four semi-structured expert interviews with senior R&D leaders across industrial, healthcare, and academic settings.
Methodological statement in the paper specifying four semi-structured expert interviews.
Because all observations come from a single practitioner, the inferential statistics are exploratory and hypothesis-generating rather than confirmatory; portability across the full portfolio awaits multi-practitioner replication.
Explicit limitation stated in the paper about the single-practitioner design and its implications for inference.
The framework is illustrated with an accounts-payable simulation and a companion spreadsheet.
Empirical illustration: the paper includes (or accompanies) an accounts-payable simulation and a spreadsheet to demonstrate the model and estimation approach.
The note starts from a compact dashboard expression, expands it into a fuller structural model, defines all variables and parameters, and shows how each cost category can be estimated from operational data.
Methodological description in the paper: construction of dashboard, expansion to structural model, full variable/parameter definitions, and stated procedures for estimating cost categories from operational data; accompanied by worked examples.
Agentic Technical Debt is a stock of accumulated design and governance liability.
Definition provided in the paper as part of the conceptual framework that labels Agentic Technical Debt as a stock (accumulated) liability tied to design and governance.
This note develops a formal and managerially usable model that distinguishes Agentic Technical Debt from Stochastic Tax.
Author states development of a formal, managerially usable model and explicit distinction between the two constructs; supported by model construction in the paper (structural model and dashboard).
Agentic AI systems combine probabilistic reasoning with delegated action through tools, context, memory, orchestration, and external workflow integration.
Conceptual/definitional statement in the paper; presented as the working characterization of 'Agentic AI systems' within the model specification.