Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Qwen3 embeddings with 300-token chunk size achieved 94.6% accuracy on a clinical question-answering benchmark.
Optimization experiment on a physician-authored clinical question-answering benchmark; best-performing configuration reported as qwen3 embeddings with 300-token chunks and 94.6% accuracy.
The system delivers sub-second query latency: median 237 ms single-user, 451 ms at 20-user concurrency.
Full-scale performance characterization reported exact median latencies for single-user and 20-user concurrency.
Multi-agent reinforcement learning has emerged as a promising approach for the combined scheduling of production and transportation tasks in decentralized factories.
Literature-context claim in the paper's introduction summarizing prior research trends and motivations for applying multi-agent RL to integrated production-transport scheduling.
Modular training represents a viable alternative in environments where a single scheduling task dominates.
Empirical findings from the paper's experiments/sensitivity analysis indicating modularly trained agents perform comparably to joint training when one scheduling task (either production or transportation) is temporally dominant.
Joint training can produce superior performance compared to the best-performing combinations of dispatching rules and modular training.
Empirical evaluation reported in the study comparing joint-training multi-agent RL against modular training and dispatching-rule baselines across simulated job-shop scheduling environments with transportation resources (sensitivity analysis over resource scarcity and temporal dominance). Specific sample size / number of scenarios not stated in the abstract.
Taken together, these insights provide theoretical clarity and practical guidance for responsible GenAI integration into creative work.
Authors' stated contribution and practical recommendations derived from the conceptual framework; no empirical evaluation of guidance effectiveness provided.
The study reinterprets process-oriented creativity theories through structural parallels with GenAI.
Conceptual reanalysis and theoretical reinterpretation based on literature synthesis (paper's theoretical contribution).
The authors propose a role-based integration model that aligns GenAI capabilities with key creative functions: idea generation, synthesis, strategic framing, and facilitation.
Presentation of a novel conceptual model / framework in the paper (theoretical design); no empirical validation or measured outcomes reported.
The paper repositions GenAI as a cognitive collaborator rather than merely a productivity tool.
Argumentative / conceptual claim supported by the proposed theoretical reframing and role-based model in the paper; no empirical testing reported.
There are structural parallels between GenAI architectures and human cognition—such as heuristic search, divergent thinking, and iterative refinement.
Conceptual mapping and theoretical comparison between GenAI architecture characteristics and cognitive/creativity constructs presented in the paper (literature synthesis / theoretical argument).
The study revisits foundational creativity theories to develop a framework for integrating GenAI into creative workflows.
Paper describes a conceptual review and theoretical synthesis of foundational creativity theories leading to a proposed integration framework; methodological (theoretical / conceptual) contribution rather than empirical validation.
Generative Artificial Intelligence (GenAI) is reshaping organisational creativity by emulating cognitive processes traditionally associated with human innovation.
Paper's theoretical argument and literature-grounded conceptual claims (conceptual analysis / literature review); no empirical sample or quantitative data reported.
There is an open opportunity to support collaborative construction where users and AI jointly develop an evolving knowledge representation.
Paper's stated research opportunity and motivation based on gaps identified in prior tools and systems (conceptual argument).
In a user study where 12 participants created slide decks, MindTrellis outperformed retrieval-only baselines in knowledge organization and cognitive load, as measured by expert ratings of content coverage and structural quality.
Controlled user study reported in the paper: N = 12 participants performing slide-deck creation tasks; outcomes assessed via expert ratings of content coverage and structural quality (comparison to retrieval-only baseline).
MindTrellis is an interactive visual system where users and AI collaboratively build a dynamic knowledge graph; users can query the graph for document-grounded information and contribute by introducing new concepts, modifying relationships, and reorganizing the hierarchy.
System design and implementation described in the paper (feature description and demonstration).
The framework, all collected signals, scoring outputs, and evaluation harness are released under CC BY 4.0.
Statement of data and code release policy in the paper.
AgentPulse surfaces deployment signal absent from benchmarks; it is a methodology, not a ground-truth ranking.
Conceptual and empirical argument in the paper supported by the analyses described (correlations with external adoption proxies and divergence from benchmark-only rankings).
The Benchmark+Sentiment sub-composite correlates with VS Code installs (ρ_s=0.44, p<0.05), reported as illustrative given that only 11 of 35 agents have non-zero installs.
Spearman correlation between Benchmark+Sentiment sub-composite and VS Code installs on the 35-agent sample, with a caveat that installs are non-zero for only 11 agents; reported correlation and p-value.
The Benchmark+Sentiment sub-composite predicts Stack Overflow question volume (ρ_s=0.49, p<0.01) in the circularity-controlled test (n=35).
Circularity-controlled Spearman correlation between Benchmark+Sentiment sub-composite and Stack Overflow question volume on 35 agents; reported correlation and p-value.
A circularity-controlled test (n=35) shows the Benchmark+Sentiment sub-composite, which contains no GitHub-derived signals, predicts external adoption proxies it does not aggregate: GitHub stars (ρ_s=0.52, p<0.01).
Circularity-controlled correlation test (Spearman) between Benchmark+Sentiment sub-composite and GitHub stars on a 35-agent sample; reported Spearman correlation and p-value.
We introduce AgentPulse, a continuous evaluation framework scoring 50 agents across 10 workload categories along four factors (Benchmark Performance, Adoption Signals, Community Sentiment, and Ecosystem Health) aggregated from 18 real-time signals across GitHub, package registries, IDE marketplaces, social platforms, and benchmark leaderboards.
Methodological description in the paper; reported sample of 50 agents and use of 18 signals from enumerated sources.
These results provide concrete tier-selection guidance across deployment scales from a single seminar to a university-wide rollout.
Concluding claim in paper based on empirical latency and cost comparisons across throughput tiers and concurrency up to 50 users from a live deployment.
Provisioned Throughput, expensive under continuous provisioning, becomes cost-competitive for institutions that can predict and concentrate their traffic toward high utilization.
Cost modeling and comparison across provisioning modes in the paper showing trade-offs conditional on utilization; based on the instrumented tiers and assumed usage patterns.
Cost analysis places both pay-per-token tiers well below the price of a STEM textbook per student per semester under a worst-case usage ceiling.
Cost analysis presented in the paper comparing per-student per-semester costs under a stated worst-case usage ceiling to the price of a STEM textbook; exact numeric assumptions not provided in the excerpt.
Priority PayGo maintains flat sub-4-second response times across the full load range.
Empirical latency measurements from the ITAS instrumentation described in the paper (over 3,000 requests across concurrency levels up to 50 and three throughput tiers).
Multi-agent LLM tutoring systems improve response quality through agent specialization.
Statement in paper describing design rationale; no quantitative quality comparison or metrics provided in the excerpt.
Generative artificial intelligence (genAI) is rapidly reshaping how knowledge and culture are produced and consumed.
Author's descriptive statement based on observed changes in production/consumption patterns (no empirical sample reported in paper abstract).
The paper provides a structured overview of energy forecasting use cases along three main dimensions: stakeholders, attributes, and data categories.
Methodological contribution described in the paper (framework/overview component).
The FETS benchmark collects and analyzes 54 datasets across 9 data categories guided by typical stakeholder interests.
Descriptive claim about the benchmark dataset compilation reported in the paper (dataset count and category count).
Foundation models show improved performance at higher aggregation levels such as national load, district heating, and power grid data.
Subset analyses within the benchmark comparing performance across aggregation levels (e.g., national load, district heating, power grid) showing better results for aggregated data.
Foundation models outperform classical machine learning approaches despite the latter having seen the full historic target data during training.
Benchmark setup where classical ML baselines had access to full historic target data during training while foundation models were pretrained/generalized; empirical comparison across the dataset collection.
Covariate-informed foundation models achieve the strongest performance.
Benchmark experiments that compare foundation model variants, including those that incorporate covariates, across the collected datasets; reported as a comparative finding in the benchmark results.
Foundation models consistently outperform dataset-specific optimized machine learning approaches across all settings and data categories.
Empirical benchmark comparing foundation models vs classical dataset-specific ML approaches across multiple forecasting settings and data categories; reported analysis uses the collection of datasets assembled for the FETS benchmark.
This paper presents the first systematic study of token consumption patterns in agentic coding tasks, analyzing trajectories from eight frontier LLMs on SWE-bench Verified and evaluating models' ability to predict their own token costs before task execution.
Stated scope and methodology in the paper: dataset is SWE-bench Verified, eight frontier LLMs were analyzed, and experiments included model self-prediction evaluation.
Reducing variability in solder-joint quality and cycle time.
Abstract statement that variability in solder-joint quality and cycle time was reduced during the deployment (no quantitative variability metrics provided in the abstract).
It maintained near-human takt time.
Abstract claim comparing the system's cycle/takt time to human performance during the deployment (no numeric takt-time comparison provided in the abstract).
Achieving a 99.4% pass rate on product-level quality-control tests.
Reported QC pass rate from the production run in the abstract (presumably based on the produced motors).
Operating without physical fencing.
Abstract statement that the run occurred "without physical fencing" (implying operation around people without traditional fences).
Produced 108 motors.
Count of products produced during the continuous run reported in the abstract.
The system operated continuously for 5 h 10 min.
Reported continuous operation duration from the production run described in the abstract.
Less than 20 min of real-world data per task.
Reported training data requirement for the deployed tasks in the authors' field experiment (abstract statement).
With less than 20 min of real-world data per task, the system operated continuously for 5 h 10 min, producing 108 motors without physical fencing and achieving a 99.4% pass rate on product-level quality-control tests.
Single field deployment / production run reported in the paper; numbers reported in the abstract (training data time, continuous operation duration, number of motors produced, fencing status, QC pass rate).
We deployed the system on an electric-motor production line to automate deformable cable insertion and soldering under real manufacturing constraints, a step previously performed manually by human workers.
Field deployment on an actual electric-motor production line described by the authors (deployment + task specification).
We present Learning-Augmented Robotic Automation, a hybrid system that integrates learned task controllers and a neural 3D safety monitor into conventional industrial workflows.
Description of the system developed by the authors (system design/development reported in the paper).
Self-correction should be treated not as a default behavior, but as a control decision governed by measurable error dynamics.
Synthesis of theoretical framing (Markov model and diagnostic inequality) and empirical results across multiple models/datasets showing thresholds and promptability of EIR.
A 'verify-first' prompt ablation on GPT-4o-mini reduces EIR from 2% to 0% and turns -6.2 pp degradation into +0.2 pp (paired McNemar p < 10^-4).
A prompt-ablation experiment reported for GPT-4o-mini showing EIR dropping from 2% to 0% and the observed accuracy change flipping from -6.2 percentage points to +0.2 percentage points; statistical significance assessed with a paired McNemar test (p < 10^-4).
In this framework, EIR functions as a stability margin and prompting functions as lightweight controller design.
Conceptual framing in the paper (cybernetic feedback loop where the same language model is controller and plant), supported by associated experiments showing prompt changes affect EIR and outcomes.
Iterate only when ECR/EIR > Acc/(1 - Acc).
The paper frames self-correction as a two-state Markov model over {Correct, Incorrect} and derives this deployment diagnostic analytically from that model.
Im Forschungskontext sind kontextbezogene Schulungs- und Begleitmaßnahmen entscheidend für den Erfolg der Copilot-Einführung.
Schlussfolgerung der Autoren aus den Befunden zur zeitlichen Entwicklung der Bewertungen wissenschaftlicher Mitarbeitender und zu unterschiedlichen Nutzenwahrnehmungen (im Abstract genannt).
Die Untersuchung zeigt, dass Microsoft 365 Copilot insbesondere im administrativen Bereich Effizienzgewinne ermöglicht.
Selbstberichtete Einschätzungen der Beschäftigten (speziell Verwaltungsmitarbeitende) in der wiederholten Querschnittsbefragung; Autoren ziehen daraus praktische Relevanz im administrativen Bereich (Abstract).