Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We introduce a novel evaluation methodology and benchmark designed to rigorously measure an LLM agent's reasoning for financial time-series analysis.
Paper describes a new methodology and benchmark (Time Series Augmented Generation, TSAG) developed by the authors for evaluating LLM reasoning on financial time-series tasks.
Effective evaluation-driven loop scaling is a central axis for advancing LLM-driven scientific discovery, and SimpleTES provides a simple yet practical framework for realizing these gains.
High-level claim supported by the aggregate experimental results and discussion in the paper.
When post-trained on successful trajectories, models not only improve efficiency on seen problems but also generalize to unseen problems, discovering solutions that base models fail to uncover.
Experiments in which models were post-trained on successful SimpleTES trajectories and evaluated on both seen and unseen problems (paper claim of improved efficiency and generalization).
SimpleTES produces trajectory-level histories that naturally supervise feedback-driven learning.
Methodological claim and supporting experiments where SimpleTES generates solution trajectories that are then used as supervision for learning.
We discovered new Erdos minimum overlap constructions that surpass the best-known results.
Reported novel combinatorial constructions (Erdos minimum overlap) in the experiments that improve on prior best-known results.
We designed quantum circuit routing policies that reduce gate overhead by 24.5%.
Experimental results reported for quantum circuit routing tasks showing a 24.5% reduction in gate overhead when using SimpleTES-designed policies.
We sped up the widely used LASSO algorithm by over 2x.
Benchmarking experiment reported in the paper comparing LASSO runtime/performance with and without SimpleTES (paper states >2x speedup).
SimpleTES consistently outperforms both frontier-model baselines and sophisticated optimization pipelines.
Comparative experimental evaluation vs. frontier-model baselines and optimization pipelines across the reported problems (paper claim).
Across 21 scientific problems spanning six domains, SimpleTES discovers state-of-the-art solutions using gpt-oss models.
Empirical experiments reported across 21 problems in six domains using gpt-oss models (paper states 21 problems).
We introduce Simple Test-time Evaluation-driven Scaling (SimpleTES), a general framework that strategically combines parallel exploration, feedback-driven refinement, and local selection.
Methodological contribution described in the paper (framework design and algorithmic description).
We propose seven interface primitives operationalizing verification-centered HCI.
Design contribution: specification of seven interface primitives within the paper (conceptual/design proposal); no user-study or empirical validation reported.
We map synthetic literacy -- oral input generating literate output -- as the defining feature of this transition.
Conceptual mapping and theoretical framing within the paper; supported by examples from technology trends but no empirical evaluation reported.
Knowledge workers become adversarial auditors rather than keystroke-producers.
Projected role-shift based on the verification-bottleneck thesis and interdisciplinary supporting arguments; no empirical longitudinal workforce study reported.
The central contribution identifies the verification bottleneck: as AI collapses production friction, the primary constraint shifts from generation to evaluation.
Theoretical argument supported by literature synthesis across multiple fields; no direct experimental quantification provided.
We contribute design guidelines for specialized AI and articulate a vision for 'ecosystem-aware' Humble AI.
Paper's stated contributions (design guidelines and conceptual vision) described in the abstract.
Qualitatively, participants used AVA as a specialized 'evidence engine'; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations.
Qualitative findings from surveys and 20 interviews reported in the paper (participant quotations and thematic analysis implied in abstract).
Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly.
Quantitative claim reported in the paper based on Difference-in-Differences analysis of usage/engagement data from the evaluation (implicit sample drawn from the >2,200 participants).
AVA operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection).
Design claim describing implemented mechanisms in the platform; described in the paper as operational features.
AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses.
System design and capability claim in the paper (description of multi-agent pipeline producing evidence-based syntheses).
AVA is a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities.
System description provided in the paper; statement of dataset size and functionality (library count and multilingual support).
Code-generating Artificial Intelligence has gained popularity within both professional and educational programming settings over the past several years.
Background statement in the paper's introduction (observational claim about recent trends in AI adoption).
The emotional effect of the human teammate was significantly more positive and arousing compared to working with Copilot.
Subjective emotion measures (valence/arousal) collected in the study; reported significant differences favoring human teammate on positivity and arousal (n=22).
Several dimensions of participants' workload were significantly reduced when using GitHub Copilot.
Subjective workload measures collected during the experiment; multiple workload dimensions reported as significantly lower in the Copilot condition (n=22).
Participants performed significantly better with GitHub Copilot than with their human teammate.
Experimental comparison of task performance between Copilot-assisted individual condition and human pair condition; statistical significance reported in results (sample size n=22).
Evaluation demonstrates speed improvements of 6-7 minutes over traditional methods.
Reported empirical timing result in paper abstract: 6-7 minutes (presumably time to validate a change) compared to traditional methods (no further detail or sample size in abstract).
Evaluation demonstrates diagnostic coverage of 92-96%.
Reported empirical range in paper abstract (92-96% diagnostic coverage over evaluated cases; specific n not provided in abstract).
Evaluation demonstrates promising results in error detection (100%).
Reported empirical result in paper abstract: 100% error detection over evaluated scenarios (no sample size given in abstract).
By orchestrating agent collaboration atop this digital twin, Aether enables automated, rapid network change validation while reducing manual effort, minimizing errors, and improving operational agility and cost-effectiveness.
High-level claim supported by system design and subsequent empirical evaluation reported in paper (evaluation details referenced in abstract).
Aether agents use a unified Network Digital Twin integrating modeling, simulation, and emulation to maintain a consistent, up-to-date network view for verification and testing.
Design claim describing the digital twin's capabilities (modeling, simulation, emulation) as part of the system; presented in paper text.
Aether features an agentic architecture with five specialized Network Operations AI agents that collaboratively handle the change validation lifecycle from intent analysis to network verification and testing.
System architecture claim in paper describing five specialized agents (design specification; no empirical sample size).
Aether integrates Generative Agentic AI with a multi-functional Network Digital Twin to automate and streamline network change validation workflows.
Paper describes Aether system design and architecture combining agentic AI and a digital twin (design-level claim; architectural description).
To mitigate the curse of dimensionality in HRL, the paper introduces a capacity-aware state–action encoding mechanism that compresses the control interface into structured summary signals.
Methodological contribution described in the paper: proposed encoding mechanism intended to reduce state-action dimensionality and simplify the control interface.
Intelligent textile technologies can effectively enhance operational efficiency in the textile industry's supply chain.
Overall result statement summarizing pilot study outcomes (inventory turnover, order fulfillment, cost control) as evidence; no numeric aggregate efficiency measure or sample size provided in the excerpt.
Intelligent textile technologies can effectively enhance supply chain transparency.
Conclusion based on the pilot study and the inclusion of blockchain-based data sharing in the model; no empirical transparency metrics or sample size reported in the provided text.
Intelligent textile technologies can effectively enhance supply chain collaboration.
Conclusion drawn from the pilot study reported in the paper; no quantitative measures of collaboration or supporting statistics provided in the supplied text.
A pilot study demonstrates significant improvements in customer satisfaction.
Reported pilot study in the paper; no details on how customer satisfaction was measured, sample size, or effect size are provided in the supplied text.
A pilot study demonstrates significant improvements in cost control.
Reported pilot study in the paper; the summary does not provide numerical cost reductions or sample size.
A pilot study demonstrates significant improvements in order fulfillment efficiency.
Reported pilot study in the paper; no sample size, quantitative metrics, or statistical tests reported in the provided text.
A pilot study demonstrates significant improvements in inventory turnover.
Reported pilot study in the paper; no sample size, numerical improvement, or statistical measures provided in the text supplied.
This study develops an intelligent supply chain model integrating AI forecasting, blockchain-based data sharing, and automated inventory management.
Methodological claim describing the authors' model development (design/integration of AI forecasting, blockchain, automated inventory management); no external validation details given here.
Intelligent textile technologies are increasingly transforming the textile industry supply chain.
Author statement in paper introduction; general trend claim (no empirical support or quantitative data reported).
The study uses a combination of cognitive systems theory, diplomatic negotiation models, and empirical Human-in-the-Loop experiments as its methodological basis.
Methods description in the paper listing theoretical foundations and empirical HITL experiments as components of the study design.
The paper outlines recommendations for international norm development, capacity building, and the creation of interoperable, transparent AI systems for diplomacy.
Policy recommendation section of the paper proposing international norms, capacity-building measures, and interoperable transparent system design.
Experimental HITL data indicate a 17% reduction in cognitive bias for hybrid human-AI teams.
Human-in-the-Loop (HITL) experiments reported in the paper; comparison of cognitive bias measures between hybrid teams and baseline (sample size not provided in summary).
Experimental HITL data indicate that hybrid human-AI teams achieved 23% faster consensus-building.
Human-in-the-Loop (HITL) experiments reported in the paper; experimental comparison between hybrid human-AI teams and baseline (details on sample size not reported in summary).
The framework is validated through real-world and simulated case studies, including UN ceasefire mediation, EU sentiment-monitoring for conflict diplomacy, and African Union peacekeeping planning.
Validation reported via a set of real-world and simulated case studies described in the paper (case study methodology; specific cases named).
Each layer augments a core dimension of diplomatic reasoning, enabling interpretable AI contributions, foresight analysis, culturally sensitive framing, and legally compliant outputs.
Conceptual mapping of each proposed layer to functional capabilities described in the paper; claimed alignment with interpretability, foresight, cultural framing, and legal compliance.
The study proposes a five-layer Human-AI collaboration architecture tailored to multilateral diplomacy consisting of: (1) Context Modeling, (2) Scenario Generation, (3) Cognitive Interfacing, (4) Decision Support, and (5) Ethical-Normative Governance.
Architectural proposal in the paper based on synthesis of literature and design choices; claimed as the output of the conceptual framework.
This paper develops the concept of Artificial Diplomacy as a structured interface between human strategic cognition and machine-supported reasoning.
Theoretical development drawing on cognitive systems theory and diplomatic negotiation models; described design and conceptual argumentation in the paper.
Policymakers can reinforce these conditions by shifting from technology-neutral principles to auditable process standards that couple AI investment with reskilling and data-quality obligations.
Policy recommendation based on the study's findings and synthesis; presented as a normative implication rather than empirically tested within the study. (Sample size not reported.)