Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The paper evaluates 'Spec Kit' and 'TDAD' as instantiations of the SGM via a four-month pilot study.
Empirical pilot evaluation reported in the paper; duration specified as four months. Sample size or number of teams/participants in pilot not specified in the summary.
The paper identifies two amplifying mechanisms for PRP: the code review bottleneck and the context window constraint.
Theoretical argumentation in the paper naming two mechanisms that amplify the PRP phenomenon (qualitative explanation).
The paper formally defines PRP with three moderating variables: task abstraction, codebase maturity, and developer experience.
Theoretical/formal definition presented in the paper identifying three moderators; claim is descriptive of the paper's conceptual model.
This paper conducted a multivocal literature review of 67 sources spanning 2022–2026.
Statement of method in the paper describing the literature review (count of sources = 67).
Telemetry across 10,000+ developers shows flat delivery metrics (no improvement in delivery outcomes) despite changes in PR and review behavior.
Observational telemetry across >10,000 developers reported in the paper; described result is no meaningful change in delivery metrics (e.g., delivery throughput, lead time) despite increases in PRs and longer reviews.
Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed.
Author assertion / methodological observation about how public benchmarks report results versus how deployments are decided; no empirical test reported in the excerpt.
A symbolic lifting operator translates simulator trajectories into qualitative descriptors, motion labels, temporal predicates, and structural diagnostics that models interpret across iterative design cycles.
Architectural/methodological contribution described in the paper: a symbolic lifting operator that converts simulator trajectories into higher-level symbolic diagnostics used by LM agents during iterative refinement.
Language model agents explore discrete topologies while numerical optimisers fit continuous parameters.
Methodological description of the system architecture in the paper (division of labor between LM agents for discrete topology search and numerical optimisers for continuous parameter fitting).
The review uses a collection of qualitative and quantitative approaches (i.e., it synthesizes both qualitative and quantitative studies).
Explicit methodological description in the abstract indicating mixed-methods literature synthesis.
A collection of qualitative and quantitative approaches reveals predictors of technological integration, including organisational preparedness, economic factors, policies, and human capital.
Statement about the review's synthesized findings from multiple qualitative and quantitative studies identifying these predictors; method = mixed-methods literature synthesis.
The primary technologies covered in this review are Electronic Health Records (EHR), telemedicine, artificial intelligence (AI), and the Internet of Things (IoT).
Explicit topical scope statement in the paper (description of review subjects); based on the paper's own selection of topics for review.
There is little empirical exploration of how professionals making high-stakes decisions perceive their agency and level of control when working with genAI systems.
Statement about a gap in the existing literature made by the authors (literature review / framing); no sample size (gap claim).
AI adoption has no detectable effects on overall employment.
Difference-in-differences estimates using administrative employment totals linked to survey-reported adoption show no statistically significant change in total employment.
As of 2024, AI adoption remains limited: about 10 per cent of firms report current use.
Newly collected firm-level survey data linked to administrative balance sheet and employer–employee records; prevalence reported in 2024 survey.
The empirical analysis uses panel data from 3,515 Chinese A-share listed firms, totaling 20,076 firm-year observations covering 2014–2022.
Statement of data and sample in the paper (sample frame and time period explicitly given).
The literature review employs the PRISMA model to screen, identify, and synthesize available literature on AI, Machine Learning and Deep Learning in promoting managerial productivity and task efficiency.
Methodological statement in the paper's abstract (explicitly states use of PRISMA for screening and synthesis).
Methodological basis: the study used analysis of aggregated industry data and a scenario approach; information sources were Russian-language materials including the Ministry of Digital Development, HSE, the Autonomous Non-Profit Organization 'Digital Economy', and analytical reviews.
Explicit methodological and data-source statements in the paper.
All artifacts associated with this study are publicly available at https://zenodo.org/records/18489222.
Statement in the paper providing a Zenodo link to artifacts.
This review identifies key research gaps and provides recommendations for future research and practice.
Authors' discussion and conclusion sections synthesizing gaps and offering recommendations based on the mapping results.
Satisfaction, Performance, and Efficiency are the most frequently investigated SPACE dimensions, whereas Communication and Activity remain underexplored.
Frequency counts and synthesis across the 39 included studies mapped to SPACE dimensions as reported by the authors.
Only 15% of the reviewed studies extend beyond three SPACE dimensions.
Authors' coding of included studies against the SPACE framework with reported proportion.
90% of the reviewed studies adopt a multi-dimensional perspective by examining at least two SPACE dimensions.
Authors' coding of included studies against the SPACE framework, yielding the reported proportion.
This paper is a systematic review and mapping of 39 peer-reviewed studies published between January 2014 and December 2024 that examine the impact of LLM-assistants on software developer productivity.
Authors conducted a systematic review and mapping exercise covering peer-reviewed studies within the stated date range; the paper reports the count of included studies as 39.
Cross-stage correlations are very weak: parsing->retrieval r = 0.14, parsing->generation r = 0.17, retrieval->generation r = 0.02.
Reported Pearson (or Spearman) correlation coefficients between stage-level metrics in the benchmark; exact correlation method not specified in excerpt.
We evaluate SecMate in a controlled study with 144 participants and 711 conversations.
Reported experimental study sample and conversation counts in the paper.
Our architecture combines a two-layer Graph Convolutional Network (GCN) encoder, twin critics, and a value network that drives the adversary.
Model architecture description in the paper specifying a 2-layer GCN encoder, twin critics, and a value network used for adversary control.
The robust backup uses the Kantorovich--Rubinstein dual, a projected subgradient inner loop, and a primal--dual risk-budget update.
Algorithmic description in the paper detailing the robust backup solver components (Kantorovich--Rubinstein dual, projected subgradient, primal-dual update).
To mitigate distributional shifts, we optimize a Soft Actor--Critic (SAC) agent against a Wasserstein-1 ambiguity set with a graph-aligned Mahalanobis ground metric that captures spatial correlations.
Methodological description of a robust training objective: SAC optimized under a Wasserstein-1 ambiguity set using a graph-aligned Mahalanobis metric to encode spatial correlations.
These intentions are projected at every decision step through a time-limited rolling mixed-integer linear program (MILP) that strictly enforces state-of-charge, port, and feeder constraints.
Method/algorithm description in the paper: a rolling MILP projection component implemented to enforce physical constraints (state-of-charge, charger port limits, feeder limits) at each decision step.
The policy learns over high-level intentions produced by a masked, temperature-annealed actor.
Method/algorithm description in the paper describing the actor design (masked, temperature-annealed) and the high-level intentions used for policy learning.
We formulate the problem as a hex-grid semi-Markov decision process (semi-MDP) with mixed actions -- discrete actions for serving, repositioning, and charging, together with continuous charging power -- and variable action durations.
Methodological description in the paper presenting the model formulation (hex-grid semi-MDP) and action space design; no external dataset required.
The analysis employs rigorous econometric methods including difference-in-differences estimation and propensity score matching to control for confounding variables across industry (NAICS 2-digit), firm size, geographic location, occupation-level characteristics, and macroeconomic conditions.
Methodological description in the paper specifying DiD and propensity score matching and listed covariates/controls.
The study uses U.S. Census Bureau Business Trends and Outlook Survey data tracking over 1.2 million businesses.
Paper statement that it incorporates the Census Bureau Business Trends and Outlook Survey covering >1,200,000 businesses.
The analysis integrates the Anthropic Economic Index capturing approximately one million AI usage interactions.
Paper statement that the Anthropic Economic Index was used and captures ~1,000,000 AI usage interactions.
Semantic search maintained comparable inter-rater agreement while reducing chart abstraction time.
Clinical utility evaluation reports that inter-rater agreement was comparable between semantic-search-assisted abstraction and clinician-performed chart review.
The authors optimized embedding model and chunking strategy using a physician-authored benchmark dataset.
Methods: experiment described as optimization of embedding model and chunking using a physician-authored benchmark dataset.
The system uses instruction-tuned qwen3-embedding-0.6B embeddings, stores vectors in a managed database with storage-optimized indexing, maintains full-text metadata in a low-latency key-value store, and operates within a HIPAA-compliant governance framework.
Methods description of system architecture and governance provided in the paper.
We deployed a semantic search system indexing 166 million clinical notes (484 million vectors) from 1.68 million patients.
Paper reports a production deployment at a large children's hospital and gives exact index counts: 166 million clinical notes, 484 million vectors, 1.68 million patients.
Through a rigorous sensitivity analysis of resource scarcity and temporal dominance, we quantify the coordination gap.
Methodological description in the paper indicating the authors performed a systematic sensitivity analysis across environmental parameters (resource scarcity and temporal dominance) to measure performance differences between training modalities.
On the n=11 subset with published SWE-bench scores, composite and benchmark-only rankings are nearly uncorrelated (ρ_s=0.25).
Spearman rank correlation between composite rankings and benchmark-only rankings on an 11-agent subset that has published SWE-bench scores; reported correlation.
We instrument ITAS, a four-agent tutoring system built on Gemini 2.5 Flash and Google Vertex AI, across three throughput tiers (Standard PayGo, Priority PayGo, and Provisioned Throughput) and eleven concurrency levels up to 50 simultaneous users, producing over 3,000 requests drawn from a live graduate STEM deployment.
Methods statement in paper describing experimental setup: four-agent ITAS built on Gemini 2.5 Flash and Google Vertex AI; three throughput tiers; eleven concurrency levels up to 50; over 3,000 requests from a live graduate STEM deployment.
The welfare consequences of genAI can be organized by a two-dimensional taxonomy: the strength of the incentive to perform the task without AI, and the severity of model collapse.
Analytical organization derived from the theoretical model presented in the paper (conceptual taxonomy based on model parameters; no empirical sample reported in abstract).
We develop a parsimonious model of behavior in collaborative interactions in which individuals can either exert human effort, rely on genAI, or refrain from work altogether.
Methodological claim: authors present a formal theoretical model with the specified choice set (model description in paper; no empirical sample reported in abstract).
Predictive performance exhibits saturation beyond a certain context length.
Experiments varying the context (input) length in foundation models and observing changes in forecasting performance; reported saturation effect in analyses.
Task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend.
Analysis comparing human expert difficulty ratings to measured token costs for tasks in SWE-bench Verified; weak alignment reported in the paper between ratings and token consumption.
Higher token usage does not translate into higher accuracy; accuracy often peaks at intermediate cost and saturates at higher costs.
Comparison of accuracy (task success) versus total token usage across runs/trajectories in the agentic coding experiments on SWE-bench Verified; reported observed relationship (peak at intermediate costs and saturation thereafter).
Learning-based control offers a more adaptive alternative, but it remains unclear whether such methods... can sustain hours of reliable operation, deliver consistent quality, and behave safely around people on a live production line.
Framing of a research gap in the paper's introduction; no primary experimental data presented here (statement of uncertainty motivating the study).
Die Studie basiert auf einer wiederholten Querschnittsbefragung lizenzierter Beschäftigter einer außeruniversitären Forschungseinrichtung.
Autorenangabe im Abstract: wiederholte Querschnittsbefragung (survey) unter lizenzieren Beschäftigten der untersuchten Forschungseinrichtung; methodische Beschreibung im Abstract.
An exploratory evaluation compared unstructured vibe coding, structured prompt engineering, and the Shift-Up approach in the development of a web application.
Paper reports an exploratory evaluation / comparative study described in the abstract; the task context is a web application development exercise comparing three approaches (no sample size reported in abstract).
This review was conducted following the guidelines of the Preferred Reporting of Items in a Systematic Review and Meta-Analysis (PRISMA).
Methodological statement in the paper's abstract indicating PRISMA adherence; no further protocol details or study counts provided in the abstract.