Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Workspace-Bench includes files up to 20GB in size.
Dataset description in the paper specifying maximum file size.
We construct Workspace-Bench with 5 worker profiles, 74 file types, 20,476 files (up to 20GB), 388 tasks, and 7,399 total rubrics, each task associated with its own file dependency graph.
Dataset construction described in the paper; counts and sizes reported by authors.
Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling them to complete both routine and advanced tasks effectively.
Conceptual definition and motivation presented by the authors in the paper (no empirical test reported).
Conceptually, AI is positioned not as an automated controller but as an intelligence-augmenting co-regulator that supports learners' capacity to coordinate effort, attention, and understanding together.
Authors' conceptual framing and interpretation of empirical results showing that adaptive and proactive AI feedback supports shared regulation processes during dyadic programming tasks.
Proactive, forecast-based feedback using machine-learning predictions of future collaboration states (Study 3) further enhances performance and sustains shared regulation by anticipating breakdowns before they manifest.
Study 3 intervention using ML-based forecasts of future collaboration states to provide proactive support; reported improvements relative to reactive/single-channel conditions (per-study sample sizes and ML model metrics not provided in abstract).
Reactive adaptive feedback (Study 2) based on real-time deviations in JME and/or JVA improves collaboration outcomes, with combined feedback targeting both dimensions yielding the strongest improvements in performance, regulatory coherence, and cognitive-to-attentional causality, outperforming single-channel feedback.
Study 2 experimental intervention delivering reactive adaptive feedback tied to real-time JME/JVA deviations; comparisons between combined feedback, single-channel feedback, and presumably control conditions reported in paper (no per-condition sample sizes in abstract).
In Study 1, there is a stable causal relationship in which JME predicts JVA (cognitive alignment drives attentional coordination).
Causal modeling / causal inference applied to time-series measures (JME, JVA) from dual eye-tracking and pupillometry in Study 1; reported directionality JME -> JVA.
High-performing dyads show a greater prevalence of productive high-JME–high-JVA episodes.
Episode-based analysis in Study 1 using eye-tracking/pupillometry to identify and count high-JME–high-JVA episodes; comparison between performance groups reported in paper.
In natural collaboration (Study 1), high-performing dyads exhibit significantly higher joint mental effort (JME) and joint visual attention (JVA) than lower-performing dyads.
Study 1 empirical comparison of dyads during collaborative debugging using dual eye-tracking and pupillometry; performance-based grouping (high-performing vs lower-performing). Exact per-study sample not specified in abstract.
Results provide operations managers with tech-backed playbooks for responsible resource use without compromising profit motives, enabling operational excellence while meeting environmental and social responsibilities.
Paper conclusion/implication statement asserting managerial applicability of findings; grounded in the study's reported results but presented as a recommendation/implication rather than a quantified finding.
Firms maintain competitive costs while implementing AI-IoT eco-networks.
Paper claims that waste and emissions reductions are achieved without compromising costs; specific cost metrics or statistical tests not provided in the abstract.
Firms embracing AI-IoT eco-networks trim carbon output by 20-35%.
Paper results reported as empirical findings; presumably measured via carbon footprint assessments and IoT/operational metrics across the case study firms and facilities.
Firms embracing AI-IoT eco-networks cut waste by 30-50%.
Paper results reported as empirical findings; based on mixed-methods case studies of 12 multinational companies and IoT data from 45 facilities (as stated in methods).
By placing networked IoT sensors in factories, trucks, storage sites, and upstream suppliers, real-time data were paired with machine-learning routines to schedule preventive maintenance, forecast orders, and guide blockchain tracking, routing adjustments, and automated decisions balancing green goals with everyday performance.
Paper description of system design and interventions: placement of sensors across supply chain nodes and pairing with ML routines for maintenance, forecasting, blockchain tracking, routing, and automated decisions.
Overall, GenAI coding assistants can increase developer productivity, although these gains depend strongly on context.
Synthesis of meta-analytic result showing a pooled positive effect on productivity (g = 0.33) together with reported substantial heterogeneity across settings and moderator analyses.
GenAI assistance produces a statistically significant, moderate positive effect on developer productivity (Hedges' g = 0.33, 95% CI [0.09, 0.58]).
Meta-analysis pooling k = 27 effect sizes (from n = 23 studies) using Hedges' g; the paper reports the pooled estimate and 95% confidence interval.
AI development may widen income disparities across industries.
Further cross-industry analysis reported in the paper indicating that AI-related development is associated with greater inter-industry income dispersion.
The effect of AI development on the firm-level skill premium is more pronounced in firms operating in industries with lower market concentration.
Heterogeneity analysis by industry market concentration (industry-level concentration measures used to stratify firms).
The effect of AI development on the firm-level skill premium is more pronounced in firms with higher levels of digitalization.
Heterogeneity analysis using measures of firm digitalization to split the sample and compare effects.
The effect of AI development on the firm-level skill premium is more pronounced in non-state-owned firms.
Heterogeneity analysis / subgroup regressions reported in the paper comparing ownership types (state-owned vs non-state-owned firms).
The main findings remain robust after addressing endogeneity using an instrumental variable approach and conducting a series of robustness checks (alternative constructions/measures, AI pilot zone policy shock tests, alternative sample restrictions).
Reported IV analysis and multiple robustness checks in the paper (alternative dependent variable constructions, alternative AI measures, policy shock tests, sample restrictions).
AI increases the firm-level skill premium by facilitating technological upgrading.
Mechanism analysis showing AI development correlates with indicators of technological upgrading or innovation within firms.
AI increases the firm-level skill premium by promoting capital deepening.
Mechanism analysis in the paper indicating AI development is associated with higher capital intensity / capital deepening at the firm level.
AI increases the firm-level skill premium by improving firm productivity.
Mechanism analysis showing positive association between AI development and measures of firm productivity in regression analyses.
AI development significantly increases the firm-level skill premium.
Econometric analysis on Chinese listed firms using the constructed firm-level AI development measure; baseline regressions reported, with endogeneity addressed using an instrumental variable (IV) approach.
Proactive feedback produces post-intervention gains in Joint Visual Attention (JVA) and Joint Mental Effort (JME).
Within-subject empirical study with 26 dyads reporting post-intervention increases in JVA and JME measures following proactive feedback.
Proactive feedback significantly improves feedback uptake.
Reported results from the within-subject study (26 dyads) indicating higher uptake/adoption of feedback when proactive feedback was provided.
Proactive feedback significantly improves task efficiency.
Within-subject empirical study with 26 dyads reported in the paper; authors report significant improvement in task efficiency for proactive feedback condition.
In a within-subject study with 26 pair-programming dyads, proactive feedback significantly improves debugging success.
Within-subject empirical study reported in the paper with 26 pair-programming dyads; statistical claim of significant improvement in debugging success under proactive feedback condition.
ProPACT uses a hierarchical adaptive policy that delivers minimally intrusive scaffolds while fading support during productive collaboration.
Algorithm/policy design described in the paper (hierarchical adaptive policy and scaffold delivery/fading behavior).
ProPACT employs an XGBoost-based forecasting model to predict emerging suboptimal collaboration states up to 30 seconds in advance.
Modeling and evaluation described in the paper; forecasting model implementation stated as XGBoost and claim of 30-second-ahead prediction (trained/evaluated on study data from the paper).
ProPACT constructs a multimodal dyadic learner model based on Joint Visual Attention (JVA), Joint Mental Effort (JME), and individual mental effort.
System design / modeling description in the paper (multimodal dyadic learner model specification).
ProPACT is a proactive AI-driven adaptive collaborative tutor that treats collaboration itself as the object of instruction.
System description presented in the paper (design/implementation claim); authors introduce ProPACT as an AI-driven adaptive collaborative tutor.
Exposing the curated index as a callable tool suggests that each frontier system can close most of the recall gap by swapping generic web search for a curated index behind the same chat interface.
Authors' argument/inference based on exposing the index as an MCP server; phrased as a suggestion (not reported as fully executed across all frontier models).
Gosset achieved perfect precision and 100% recall against the cross-system union of verified drugs.
Empirical evaluation reported in paper using the cross-system union of verified drugs as the reference set across the 10-target benchmark.
Across 10 targets Gosset returns 3.2x more verified drugs per query than the best frontier system.
Empirical result from the authors' benchmark comparing verified drugs-per-query across systems on 10 targets.
General-purpose LLMs with web search are increasingly used to scout the competitive landscape of pharmaceutical pipelines.
Statement in paper's introduction/background; no empirical measurement provided in the excerpt.
These evolved models improve downstream end-to-end agentic data-science (ADS) performance, increasing performance for Copilot CLI, Claude Code, and Codex on the BLADE benchmark by up to 73%.
Empirical evaluation on the BLADE benchmark comparing downstream ADS performance using Copilot CLI, Claude Code, and Codex with and without the evolved models; reported maximum improvement 'up to 73%'.
The evolved models generalize to new datasets.
Reported experiments showing performance of the evolved models on datasets not used during evolution / training (as described in the paper's experimental results).
The evolved models jointly improve agent-facing interpretability (as measured by the LLM-based metric) and generalize to new interpretability tests.
Experimental evaluation using the proposed LLM-based interpretability metric, including tests on held-out interpretability evaluations described in the paper.
The evolved models jointly improve predictive performance.
Experimental results reported in the paper comparing evolved models to baselines on predictive metrics across datasets (details of datasets and metrics referenced in the experiments section).
We introduce a novel LLM-based interpretability metric that measures a suite of LLM-graded tests probing whether a fitted model's string representation is 'simulatable' by an LLM (i.e., whether the LLM can answer questions about the model's behavior by reading its string output alone).
Design and specification of an interpretability metric based on LLM-graded tests, described in the paper; metric operationalized by asking LLMs questions about models' string representations.
Agentic-imodels develops a library of scikit-learn-compatible regressors for tabular data that are optimized for both predictive performance and a novel LLM-based interpretability metric.
Implemented library of scikit-learn-compatible regressors described in the paper and used in experiments; optimization objective includes predictive performance and an LLM-based interpretability metric.
We introduce Agentic-imodels, an agentic autoresearch loop that evolves data-science tools designed to be interpretable by agents.
Description and implementation of a new system (Agentic-imodels) presented in the paper; methodological contribution described as an autoresearch loop that evolves tools.
The United States' existing public active labor market programming (WIOA) can support baseline wage recovery for vulnerable populations.
Aggregate results from the WIOA records (2017-2023) indicate general wage recovery among participants, interpreted as baseline support for vulnerable populations.
Employer-led programs—most notably apprenticeships—are associated with the highest incidence of successful outcomes.
Comparative analysis of program types within the WIOA dataset (2017-2023) showing employer-led interventions (apprenticeships) have higher rates on the Retrainability Index / success metrics than other program types.
Successful WIOA outcomes are driven mostly by post-program wage gains (possibly due to 'catch-up' mean reversion) rather than by occupational changes.
Decomposition of the Retrainability Index on the WIOA dataset (2017-2023) shows that observed program 'success' corresponds primarily to wage recovery measures rather than large shifts in RTI/occupation; authors note mean reversion as a possible explanation.
Mechanism tests reveal efficiency gains via automation are a key pathway by which AI increases productivity in constrained firms.
Mechanism analysis reported in the paper (tests linking AI adoption to automation-related efficiency improvements in constrained firm clusters).
Firms constrained by limited intangibles, outdated hardware, or weak human capital benefit most from AI adoption when AI mitigates bottlenecks (i.e., larger positive TFP effects for resource-constrained firms).
Subgroup/cluster-specific estimates from panel analysis showing larger productivity gains in clusters characterized by limited intangibles, outdated hardware, or weak human capital.
AI integration reduces regulatory breaches.
Reported empirical association showing fewer regulatory breaches following AI integration, estimated via System GMM with FE and RE robustness checks (no count data or effect sizes provided in the supplied summary).