Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Applying DeePC yields measurable improvements in system-level outcomes (reduced total travel time and CO2 emissions) in a very large, high-fidelity microscopic simulation of Zürich.
Simulation experiments in a city-scale, high-fidelity microscopic closed-loop simulator of Zürich comparing DeePC-controlled signals against baseline controllers (e.g., fixed-time or standard adaptive schemes); reported reductions in aggregated metrics (total travel time and CO2 emissions).
A model-free traffic control approach (DeePC) can steer urban traffic via dynamic traffic-light control without building explicit traffic models.
Algorithmic/theoretical development (behavioral systems theory + DeePC) and controller-in-loop experiments in a high-fidelity microscopic closed-loop simulator of Zürich demonstrating closed-loop control using only input–output trajectory data (Hankel matrices) rather than parametric model identification.
Traditional machine-learning baselines were included for comparison in the benchmarks.
Paper explicitly states that traditional ML baselines were used alongside TSFMs in benchmarking experiments. The summary does not list which baselines or their quantitative results.
The dataset sampling resolution is at the millisecond level, enabling forecasting horizons from 1 step (100 ms) up to 96 steps (9.6 s).
Paper states sampling resolution is millisecond-level and defines forecasting tasks spanning 1 to 96 steps (100 ms to 9.6 s). This is a methodological description rather than an experimental metric.
Introduces a new millisecond-resolution dataset of wireless channel and traffic-condition measurements from an operational 5G deployment.
Paper describes collection of operational 5G telemetry at millisecond sampling resolution; dataset is presented as a novel domain addition to TSFM pretraining corpora. Exact number of records/sessions not specified in the provided summary.
Historical transitions in standard work hours (e.g., six-day to five-day week) show that phased implementation, collective bargaining, and complementary policies can make work-time reductions feasible and economically beneficial.
Historical analyses and case studies of past industrialized-country workweek transitions cited in the synthesis; evidence drawn from historical institutional records and prior economic histories rather than a unified econometric analysis.
The evaluation compared models on multiple metrics (accuracy, precision, recall, F1, AUC) across repeated trials and cross-company tests, and reported gains for AI methods across these metrics.
Evaluation protocol described: repeated trials, cross-validation, holdout sets, cross-company tests; reported performance improvements for AI models on the listed metrics.
Ensemble methods and deep learning models show the largest and most consistent improvements in predictive performance relative to classic statistical models.
Aggregate results across repeated trials and evaluation metrics indicate Random Forests and Gradient Boosting (ensembles) and deep neural networks outperform linear/logistic regression and other baselines on the publicly available datasets used.
Modern AI-driven prediction methods (especially ensemble models and deep neural networks) systematically outperform traditional statistical approaches at predicting job performance in publicly available workforce datasets.
Direct model comparison reported in the paper: baseline statistical models (linear/logistic regression) versus machine learning models (Random Forest, Gradient Boosting, SVM, deep neural networks) evaluated on multiple publicly available workforce datasets using cross-validation and holdout sets; performance reported on accuracy, precision, recall, F1, and AUC across repeated trials.
Research priorities include rigorous real-world trials assessing patient outcomes, cost-effectiveness, and labor impacts; comparative studies of integration strategies; measurement of long-run workforce effects; and development of standard metrics and monitoring frameworks.
Explicit recommendations from the narrative review based on identified gaps: scarcity of RCTs, economic analyses, and long-term workforce studies.
Economists and researchers should measure organizational mediators (governance, mentoring practices, learning processes) alongside AI adoption and use empirical designs such as difference-in-differences with phased rollouts, randomized mentoring/training interventions, matched employer–employee panels, and IV exploiting exogenous shocks to innovation backing to identify causal effects.
Methodological recommendations and proposed empirical designs contained in the paper; no implementation or empirical results reported.
The integrated framework links multi-level outcomes: micro (individual skills, task performance), meso (team coordination, workflows), and macro (organizational strategy, innovation, productivity) effects to adaptive structuration processes and affordance actualization.
Framework specification and theoretical mapping across levels in the conceptual paper; no empirical validation or sample.
The paper develops a conceptual framework that integrates Adaptive Structuration Theory (AST) and Affordance Actualization Theory (AAT) to explain how effective human–AI collaboration can be structured within organizations.
Conceptual/theoretical synthesis and literature integration combining AST and AAT streams; no original empirical data or sample reported (theoretical development).
As the competition progressed, teams relied more on the AI for larger subtasks (increasing delegation and reliance).
Time-series instrumentation of AI interactions and participant behavior during the live CTF with 41 participants showing increased frequency and scope of delegated tasks later in the event.
One autonomous agent finished second overall on the fresh challenge set.
Final ranking/scoreboard from benchmarking the four autonomous agents against the live CTF challenge set and human teams; agent achieved overall 2nd place.
In a live onsite Capture-the-Flag (CTF) study (41 participants), human teams increasingly delegated larger subtasks to an instrumented AI as the competition progressed.
Empirical observation and instrumentation of AI interactions during a live, onsite CTF with 41 human participants/teams; delegation and task-size metrics tracked over time during the event.
Reward shaping at the assignment layer enables an explicit trade-off between diagnostic accuracy and human labor by incorporating penalties for human involvement.
Methodology section describing reward shaping and experimental comparisons showing different accuracy/human-effort trade-offs (results reported in paper; exact experimental details not provided in the summary).
Masked reinforcement learning techniques constrain or mask action spaces, reducing exploration over huge symptom/action spaces.
Paper describes use of masked RL to limit action options during training and execution; used in both assignment and execution layers (methodological claim supported by algorithmic description and experiments).
The upper layer ('master') learns turn-by-turn human–machine assignment using masked reinforcement learning with reward shaping to balance accuracy and human cost.
Methodological description in the paper and empirical results from experiments using masked RL and reward-shaped objectives at the assignment layer (implementation and experimental setup reported; dataset/sample size not specified in summary).
Service empathy mediates the relationship between employee emotion and collaboration proficiency.
Mediation analysis conducted on the experimental sample (n = 861) showing that measured 'service empathy' accounts for (part of) the effect of employee emotion on collaboration proficiency.
The paper advances augmentation debates by articulating the leader’s practical role when decision lead‑agency shifts between humans and AI and by detailing systemic HR changes needed to sustain performance, legitimacy and well‑being.
Stated contribution of the conceptual synthesis comparing existing augmentation and leadership literatures and providing an HR‑focused framework; descriptive of the paper's intellectual contribution.
Core practice 4 — Embed governance: make accountability, bias testing, privacy safeguards, audit trails, escalation thresholds and human oversight explicit and routine.
Prescriptive governance practice grounded in literature on algorithmic accountability and risk management and in practitioner examples; presented without original empirical validation.
Core practice 3 — Manage the human–AI relationship: build adoption, psychological safety and calibrated trust; address automation anxiety and misuse.
Framework recommendation synthesizing organizational‑psychology and technology adoption literature plus practitioner observations; not tested empirically in the paper.
Core practice 2 — Treat AI outputs as hypotheses: require human sensemaking and validation rather than blind adoption of model outputs.
Prescriptive practice derived from reviewed research and practitioner cases emphasizing human oversight; presented as framework guidance rather than empirically validated intervention.
Core practice 1 — Allocate work by comparative advantage: assign tasks to humans or AI based on relative strengths (e.g., speed, pattern detection, contextual judgement).
Conceptual component of the framework drawn from synthesis of empirical findings in prior human–AI and task allocation literature and practitioner examples; no new empirical testing in the paper.
AI methods have improved molecular property prediction, protein structure modelling, ADME/Tox prediction, NLP-based extraction from literature, virtual screening, and generative chemistry, accelerating early-stage tasks.
Compilation of benchmarking results, method-comparison studies, and applied case studies cited in the paper across these specific application areas.
AI has materially improved efficiency, decision-making, and early-stage productivity in drug discovery, especially in hit discovery, property prediction, and protein modelling.
Synthesis of published benchmarking studies and industry case studies reported in the paper (e.g., improvements in virtual screening throughput, property-prediction benchmarks, and protein-structure prediction results such as those from folding competitions and tool evaluations).
Molecule operates a marketplace for decentralized clinical and preclinical assets, focusing on tokenizing drug assets and enabling investors to finance development.
Case-study description based on Molecule's public materials and marketplace listings; demonstrates platform design and transactions rather than long-term outcomes.
VitaDAO is a community-driven organization funding and acquiring IP for longevity-related research, emphasizing open science and community governance.
Detailed case-study description drawing on VitaDAO's public documentation, governance records, and whitepaper materials.
Seed 2.0 Lite achieved 75.7% success rate with-skill, an increase of +18.9 percentage points over baseline.
Model-specific reported result in the paper: Seed 2.0 Lite with-skill success rate (75.7%) and reported improvement (+18.9pp); reported from the benchmark runs.
GLM-5 Turbo achieved 78.4% success rate with-skill, an increase of +5.4 percentage points over baseline.
Model-specific reported result in the paper: GLM-5 Turbo with-skill success rate (78.4%) and reported improvement (+5.4pp); based on the benchmark evaluation.
Nemotron 120B achieved 78.4% success rate with-skill, an increase of +18.9 percentage points over baseline.
Model-specific reported result in the paper: Nemotron 120B with-skill success rate (78.4%) and reported improvement (+18.9pp); results drawn from the benchmark runs.
MiniMax M2.5 achieved 81.1% success rate with-skill, an increase of +13.5 percentage points over baseline.
Model-specific reported result in the paper: MiniMax M2.5 with-skill success rate (81.1%) and reported improvement (+13.5pp); based on subset of the 185 scenario-runs across the evaluated models.
Results across 5 open-weight model conditions and 185 scenario-runs show consistent skill lift across all models.
Aggregate experimental results reported in the paper: evaluation over 5 model conditions and 185 scenario-runs, with cross-model improvement when SKILL is provided.
AI-adopting firms increase R&D expenditures following adoption.
Firm financial data showing higher R&D spending for adopters relative to nonadopters in post-adoption periods using the diff-in-diff framework.
Post-adoption patents by AI adopters receive more citations than those of nonadopters.
Difference-in-differences estimates comparing citation counts per patent before and after AI installation versus nonadopters; patent citation data used as the dependent variable.
Firms that adopt AI subsequently increase patenting relative to nonadopters.
Firm-level analysis using a novel AI adoption measure based on timing of AI product installations and a stacked difference-in-differences design exploiting staggered adoption; dependent variable = firm patent counts (patenting rate). (Sample size and exact time period not specified in the provided text.)
Using distributed systems as a principled foundation is a useful approach for creating and evaluating LLM teams.
Primary methodological proposal of the paper; supported by conceptual argument and (per the paper) mappings between distributed-systems concepts and LLM team design (specific experimental validation not detailed in the excerpt).
Large language models (LLMs) are growing increasingly capable.
Statement in the paper's introduction/abstract summarizing the field; based on observed progress in LLM development cited by the authors (no experimental sample size provided in the excerpt).
Only seven specialized skills produce meaningful gains (up to +30%).
Empirical results showing that 7 out of 49 skills yielded meaningful positive improvements in acceptance-test pass rates, with gains up to 30%.
The average gain from injecting skills is only +1.2% in pass rate.
Aggregated pass-rate differences computed across the benchmark tasks comparing with-skill vs without-skill conditions, reported as an average +1.2% gain.
Analysis of benchmark data (n = 667) reveals substantial synergy effects: Llama-3.1-8B improves human performance by 23 percentage points.
Empirical analysis of the same benchmark dataset (n = 667) using the Bayesian IRT model; reported improvement in human performance with Llama-3.1-8B assistance of +23 percentage points.
Analysis of benchmark data (n = 667) reveals substantial synergy effects: GPT-4o improves human performance by 29 percentage points.
Empirical analysis of a benchmark dataset of n = 667 using the paper's Bayesian IRT framework; reported improvement in human performance with GPT-4o assistance of +29 percentage points.
The work offers a blueprint for converting the ideological potential of AI into implementable, regulator-compatible utilities in pharmaceutical science by synthesizing quantitative measures and practical measures.
Claim about the paper's contribution (blueprint). It is an author claim about the synthesis and guidance provided; the excerpt does not include empirical validation that following the blueprint yields successful implementation.
The paper proposes a systematized framework of integration that emphasizes creating high-impact pilot projects, in-the-wild testing, and ongoing monitoring of models in accordance with FDA, EMA, and EU AI Act guidance.
Described as the paper's proposed framework and recommendations for regulatory-aligned implementation. The excerpt indicates the proposal but does not present validation or empirical testing of the framework.
Grounded in the Resource-Based View (RBV), AI is conceptualized as a strategic intangible resource that can confer a competitive advantage when integrated with complementary capabilities.
Theoretical framing presented in the paper (RBV-based conceptualization); not an empirical finding but an explicit conceptual claim.
Firms with high AI adoption had an average profit growth rate of 9.5%, compared to 5.8% for low adopters.
Reported profit growth rates for high vs. low AI adoption groups from the questionnaire data (N=400); the paper gives the specific averages: 9.5% (high adopters) vs. 5.8% (low adopters).
O artigo discute implicações gerenciais e de políticas públicas para reduzir fricção, acelerar adoção responsável e orientar investimentos em produtividade e inclusão.
Seção de discussão mencionada no resumo abordando encargos gerenciais e políticas públicas; não há avaliação empírica de políticas no resumo.
O artigo entrega instrumentos replicáveis — a escala SCF-30, um checklist de governança mínima de IA e uma matriz 30-60-90 dias — para uso prático.
Afirmação explícita no resumo de que instrumentos replicáveis são disponibilizados; presunção de inclusão dos instrumentos no corpo do artigo.
AI significantly enhances firms' total factor productivity (TFP).
Empirical results from the multidimensional fixed-effects panel model applied to the 2007–2023 sample of agricultural A-share firms; statistical significance reported in the paper.