Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Per-criterion accuracy climbs with stronger models.
Empirical comparison across model strengths reported in the Harvey LAB study (12,510 trajectories) showing per-criterion accuracy trends correlated with model strength.
The findings imply that research evaluation and science policy should adopt assessment frameworks that distinguish between recombinant and conceptual forms of creativity and recognize that different modes of AI adoption produce different types of scientific contribution.
Policy/recommendation statement grounded in the paper's empirical findings on heterogeneous creativity effects by AI research mode.
Adaptation-oriented AI research (modifying AI models for domain-specific problems) is associated with relatively higher object-based creativity.
Subgroup/heterogeneity analysis in the OpenAlex dataset classifying AI publications by research mode (Adaptation-oriented) and comparing object novelty outcomes across modes.
Tool-oriented AI research (applying existing AI models to domain tasks) is associated with the largest gains in recombinant-based creativity.
Subgroup/heterogeneity analysis in the OpenAlex dataset classifying AI publications by research mode (Tool-oriented) and comparing recombinant novelty outcomes across modes.
AI publications have a 5.5 to 10.2 percentage point higher likelihood to rank in the top creativity decile.
Reported quantitative effect from the paper comparing top-decile creativity probabilities between AI and non-AI publications in the OpenAlex sample.
AI publications are significantly more likely to achieve top-decile creativity relative to non-AI publications.
Observational statistical analysis comparing AI-labeled vs non-AI publications across novelty and impact measures using the >1M OpenAlex dataset (novelty measured as recombinant and object novelty; impact measured as 3-year and 10-year citation impact).
Network composition analysis of 8,012 workers shows all have inference-capable hardware.
Network composition analysis covering 8,012 workers; hardware capability inferred from worker-reported or probed specifications.
Experts assigned the highest responsibility for addressing these risks to general-purpose AI developers and governance actors (including governments, regulators, and standards bodies).
Delphi ratings of actor responsibility reported in paper: highest responsibility attributed to general-purpose AI developers and governance actors by 272 experts.
Policymakers in emerging economies should adopt integrated policy frameworks combining AI development incentives, labour market reform, and education strategies to ensure technological progress translates into inclusive and sustainable development.
Policy recommendation derived from the study's empirical findings and interpretation.
The empirical findings validate the core theoretical proposition of Routine-Biased Technological Change that skill-biased technological change operates through heterogeneous channels invisible at the aggregate level.
Synthesis of empirical results (skill-disaggregated effects differ, total unemployment insignificant) used to support the RBTC theoretical proposition.
Unemployment among less-educated workers shows a positive long-run relationship with sustainable development, interpreted as reflecting structural labour reallocation effects consistent with RBTC.
Long-run ARDL coefficient for less-educated workers' unemployment reported as positive in the paper; interpretive link to Routine-Biased Technological Change (RBTC) and labour reallocation.
In the long run, AI adoption contributes positively and significantly to sustainable development through productivity gains and innovation spillovers after structural adjustments are completed.
Long-run ARDL estimates reported in the paper indicating a positive and statistically significant long-run coefficient for AI adoption; theoretical interpretation invoking productivity gains and innovation spillovers.
We will open-source all evaluation codes, tasks, and data at https://github.com/mrwwk/DeskCraft.
Author statement promising release of code, tasks, and data (stated in abstract).
GPT-5.4 reaches 27.6% on interactive tasks.
Author-reported benchmark result for GPT-5.4 on interactive tasks from the evaluation (reported in abstract); presumably measured across the evaluation tasks.
GPT-5.4 reaches 31.6% on standard tasks.
Author-reported benchmark result for GPT-5.4 on standard tasks from the evaluation (reported in abstract); presumably measured across the evaluation tasks.
We evaluate 18 proprietary and open source agents on 538 tasks.
Author-reported evaluation methodology and scale (number of agents and tasks) as stated in abstract.
Mid-turn interaction captures both agent-initiated clarification under uncertainty and user-initiated interruption during execution, while post-turn interaction accommodates user-driven feedback after the agent signals completion.
Author description of interaction protocol structure (design specification in paper abstract).
DeskCraft formalizes human-agent collaboration into an interaction protocol covering mid-turn and post-turn exchanges.
Author statement in abstract describing the protocol (design/method contribution).
DeskCraft covers professional creative software across design, video, audio, and 3D creation.
Author statement in abstract listing covered software domains.
DeskCraft organizes tasks into a multilevel difficulty taxonomy, with long horizon tasks requiring over 50 execution steps.
Benchmark design described in abstract (explicit statement that long-horizon tasks require over 50 execution steps).
We introduce DeskCraft, a desktop GUI benchmark targeting long horizon creative and engineering workflows and proactive human-agent collaboration.
Author statement describing the new benchmark (benchmark design and scope described in paper).
The paper constructs firm-level indicators of artificial intelligence and new quality productive forces for new energy vehicle firms.
Authors state they constructed firm-level indicators as part of their empirical approach on the Yangtze River Delta panel dataset.
Artificial intelligence affects firms' new quality productive forces through improvement of innovation output.
Mechanism tests reported by the authors showing empirical evidence that AI improves innovation output (e.g., measured innovation outcomes) which is linked to higher new quality productive forces.
Artificial intelligence affects firms' new quality productive forces through optimization of R&D personnel structure.
Mechanism tests reported by the authors using the constructed indicators and panel data; empirical evidence cited that links AI to changes in R&D personnel structure which in turn link to new quality productive forces.
The promoting effect of artificial intelligence on new quality productive forces is more pronounced among small-sized enterprises.
Heterogeneity tests by firm size in the panel data; authors report stronger positive effects for small-sized firms.
The promoting effect of artificial intelligence on new quality productive forces is more pronounced in Jiangsu and Zhejiang provinces.
Heterogeneity tests on the Yangtze River Delta panel data comparing regional subsamples; authors report stronger positive effects in Jiangsu and Zhejiang.
The positive effect of artificial intelligence on firms' new quality productive forces remains robust after addressing endogeneity concerns and conducting robustness checks.
Authors report endogeneity-corrected estimations and multiple robustness checks on the same panel dataset and constructed firm-level indicators; specific endogeneity correction methods and robustness checks are not detailed in the excerpt.
Artificial intelligence significantly promotes the growth of new quality productive forces in new energy vehicle firms.
Panel data analysis of new energy vehicle firms in the Yangtze River Delta from 2001 to 2023; firm-level indicators of artificial intelligence and new quality productive forces constructed; regression estimation showing a significant positive effect.
Proactive, edge-side prompt optimization can substantially reduce inference costs without sacrificing coding quality.
Aggregate experimental results on token reductions and preserved/improved task accuracy reported in the paper.
Compared with LLMLingua-2 at matched compression rates, our method consistently achieves superior OckScore performance across all evaluated backends.
Head-to-head experimental comparison reported in the paper between the proposed middleware and LLMLingua-2 (matched compression rates) measuring OckScore.
Ablation studies indicate that the gains come primarily from the structural rewriting stage rather than simple function-name extraction.
Ablation experiments reported in the paper comparing full rewrite pipeline versus variants (e.g., function-name extraction only).
Prompt compression via the middleware preserves or improves task accuracy on the evaluated benchmark.
Reported task accuracy comparisons on OMH-Polyglot before and after applying middleware across evaluated backends.
The middleware reduces total tokens (prompt + completion) by up to 18.8 percent.
Empirical measurements reported in the paper comparing total token usage (prompt + completion) with and without middleware.
Across three commercial LLM backends, the middleware reduces prompt tokens by 34–47 percent.
Empirical results reported from experiments on OMH-Polyglot across three commercial LLM backends (aggregate token counts before vs. after middleware).
We introduce a pre-flight, edge-side prompt-rewriting middleware that runs locally (using Llama 3.2 (3B)) to perform cross-lingual translation into English, structural rewriting into a compact task-oriented format, and regex-validated rewrite-with-fallback safeguards to ensure the optimized prompt is never larger than the original.
System implementation and design described in the paper (local Llama 3.2 (3B) model, translation, rewriting, and rewrite-with-fallback mechanism).
The positive influence of industrial robot application on MVCR is especially significant in low-technology industries.
Heterogeneity/subsample analysis reported in the paper showing larger estimated effects in low-tech industry groups.
The positive influence of industrial robot application on MVCR is especially significant in downstream segments of the value chain.
Heterogeneity/subsample analysis reported in the paper showing larger estimated effects in downstream value-chain segments.
The positive influence of industrial robot application on MVCR is particularly significant in privately owned businesses.
Heterogeneity/subsample analysis reported in the paper showing stronger estimated effects for privately owned firms compared with other ownership types.
Industrial robot application positively impacts manufacturing value chain resilience (MVCR).
Empirical assessment using the constructed industrial-robot application indices and MVCR index (regression/empirical analysis on Chinese A-share listed firms).
The machines are increasingly becoming competent.
Authorial assertion about the trend in AI capability (no metrics or studies provided in the excerpt).
The concept of co-intelligence describes a new cognitive ecology where the human and artificial minds mutually influence one another to come up with ways of comprehending, creating and making choices that neither of them could accomplish individually.
Conceptual claim attributed to Ethan Mollick (2024) and extended by the author — described conceptually rather than demonstrated empirically in the excerpt.
None of the past technologies have spread into so many aspects of human life, so fast.
Author's comparative assertion about the speed and breadth of AI diffusion relative to prior technologies (no empirical comparison provided in the excerpt).
Artificial intelligence has become a partner in our everyday activities: it dictates our emails, diagnoses our diseases, educates our young children, controls our budgets, creates our artworks, and influences the policies made by governments and corporations.
Authorial assertion listing domains of current AI use (no empirical study or quantified data provided in the excerpt).
The internet had to cope with more or less a decade before it could reach one billion users; social media did it in half times.
Comparative historical adoption claim presented by the author (no citation or empirical method given in the excerpt).
Less than a year after its debut, hundreds of millions of individuals on all seven continents were using large language models, in virtually every field of professional activity, and in most languages.
Authorial assertion summarizing global LLM adoption (no specific study, dataset, or methodology provided in the excerpt).
There were now a hundred million ChatGPT users in two months.
Authorial assertion in the text citing a user-count milestone for ChatGPT (no study or data source provided in the excerpt).
The scientific results converged in both runs.
Paper statement reporting that the scientific results from both agents converged across the two experimental runs (descriptive outcome of the runs).
SLMs can effectively benefit from LLM assistance.
Empirical finding stated in the abstract based on experiments with hybrid MAS architectures where SLMs were assisted by cloud LLMs.
We adapt two representative MAS architectures to support hybrid inference and study how individual design choices shift the operating point along the Pareto frontier of power, cost, and performance.
Methodological statement from the paper: the authors implemented/adapted two MAS architectures and performed empirical evaluation to map Pareto trade-offs.
Hybrid multi-agent systems (MASs) combining on-device and cloud models offer a promising middle ground between LLMs and SLMs.
Framed as a motivation in the paper; the authors adapt two MAS architectures and study hybrids empirically to support this position.