Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Human-centric design principles significantly improve implementation outcomes and organizational value creation.
Systematic analysis of recent literature and case studies reported in the paper (literature review + case examples). No sample size or quantitative meta-analysis reported in the summary.
The Bayesian optimization agents obtain higher payoffs than the evaluated LLM agents on this spatial search task.
Comparative computational results reported in the paper showing Bayesian optimization agents outperform the evaluated LLM agents on the spatial search task.
In this experiment, adding a one-sentence first-round randomization instruction improves collective payoff by more than three times the estimated payoff difference across the eight network topologies.
Quantitative result from the paper's computational experiment comparing collective payoffs with and without a one-sentence first-round randomization instruction; phrased as an improvement 'more than three times' the estimated payoff difference across topologies.
The LLM agents show a significant network-efficiency effect when instructed to randomize their first-round choices, but not under the default initialization.
Computational experiments reported in the paper comparing LLM agent performance across network topologies under two initialization conditions (default vs. explicit first-round randomization). The result is framed as significant in the experiments.
The Mason--Watts experiment (PNAS 2012) showed that human groups in shorter-path networks outperform those in longer-path networks on a two-dimensional search task.
Citation of prior published experiment (Mason & Watts, PNAS 2012) reporting human-subject experimental results on an 2D spatial search task comparing network topologies.
Four implications for governance follow from this perspective: promise restraint, evidence of deployment, documentation obligations, and greater infrastructural pluralization.
Prescriptive recommendations derived from the paper's legitimacy-gap framework and its analysis of historical and contemporary dynamics.
The governance gap operates as a meta-condition of legal, moral, and political authorization for AI.
Theoretical claim within the model, based on literature in legitimacy studies and governance discussed in the paper.
The institutional assessment and commercialization gaps mediate whether AI promises are credited, funded, and productized.
Conceptual/analytic claim tied to the hierarchical model; supported by historical/literature examples in the paper.
The capability gap functions as the technical substrate of AI promises.
Conceptual claim within the proposed model; grounded in historical examples and literature synthesis presented in the paper.
A hierarchical interaction model of four legitimacy gaps explains AI development dynamics: capability gap, institutional assessment gap, commercialization gap, and governance gap.
The paper proposes a conceptual/hierarchical model based on synthesis of historical cases and literature in legitimacy studies and sociology of expectations.
These results provide preliminary evidence that enhanced structural clarity, action cues, evidence signals, and temporal validity indicators can substantially improve the reliability and efficiency of AI browser agents.
Aggregate interpretation of experimental results (comparative improvements in PASS rates, reduced PARTIALs, and lower step counts across 300 runs).
The agent-ready website lowered the average step count from 9.31 to 6.49.
Reported average step counts per run in the experiment for baseline vs agent-ready sites.
The agent-ready website reduced PARTIAL outcomes from 43 to 3 (compared to the baseline).
Reported experimental counts of PARTIAL outcomes for the two website variants (presumably out of 150 runs each).
The agent-ready website achieved 134 PASS runs out of 150 versus 74 out of 150 for the baseline (strict success rates of 89.3% vs. 49.3%).
Reported experimental results: PASS counts and computed strict success rates for each website condition (150 runs per condition).
The agent-ready framework is structured around three dimensions: agent interpretability, agent executability, and agent decision reliability, supported by features such as machine readability, semantic clarity, agent actionability, and contextual decision-reliability signals.
Framework specification and definitions presented in the paper (conceptual).
This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents.
Paper contribution: authors propose and describe a new design framework (conceptual / methodological description).
The Lab of the Future achieves this compression by tightly integrating domain‑specific computational models (Scientific AI), automated robotic experimentation (Physical AI), and intelligent orchestration (Agentic System).
Conceptual/mechanistic claim presented by the authors in the paper; no empirical validation, controlled experiments, or sample described in the excerpt.
The Lab of the Future offers a way to compress this cycle from months to days.
Author claim in the paper describing anticipated impact of the Lab of the Future; no supporting empirical trial, sample size, or quantified study provided in the excerpt.
These results provide strong evidence that AI performance on this class of work is already high and rapidly improving.
Synthesis/interpretation by the authors based on BusinessCaseBench evaluation results (aggregate claim; numeric strength of evidence not included in the excerpt).
On BusinessCaseBench, frontier AI models already score highly against instructor rubrics.
Empirical evaluation reported in the paper comparing frontier AI model outputs to instructor-derived grading rubrics on the BusinessCaseBench (specific model names, metrics, and numeric scores not provided in the excerpt).
The 'case method' used by top business schools provides a natural foundation for addressing this measurement gap.
Argument in the paper proposing the case-method pedagogy as an appropriate basis for constructing benchmarks of analytical professional work (methodological rationale; no empirical validation given in the excerpt).
Large language models (LLMs) are improving rapidly as reflected in benchmark scores.
Statement in the paper referencing trends in benchmark scores (no specific benchmark names or numeric trends provided in the excerpt).
After the 2022 LLM shock, more-exposed firms raised labor productivity (β2 = +0.075).
Estimated coefficient from the continuous-treatment DiD model on the full panel (author reports β2 = +0.075 for labor productivity).
The IT-BPM industry holds a disproportionate share of jobs highly exposed to large language models (LLMs).
Author statement based on occupational LLM-exposure scores (Eloundou et al. 2024) aggregated to industry/business-line employment shares.
The IT-BPM industry employs roughly 5.4 million workers in back-office, coding, and software roles.
Author-provided employment figure for the IT-BPM sector (stated in the paper).
The paper bridges human factors, AI interpretability, and decision science to design trustworthy, reliable, and human-centred AI teammates.
Stated intellectual/conceptual contribution: an interdisciplinary synthesis presented in the paper combining human factors, explainability, and decision science; presented as the paper's contribution rather than validated by extensive external data in the provided text.
We propose evaluation metrics for measuring collaboration effectiveness in high-stakes domains.
Methodological/conceptual contribution in the paper that defines or proposes metrics to evaluate human-AI collaboration effectiveness for high-stakes applications; this is a proposed set of metrics rather than empirically validated measures.
Adaptive explanations enhance overall team performance.
Experimental evaluation with a medical decision-support prototype reports improved team-level performance metrics when adaptive explanations are employed; details on team size, number of teams, and statistical measures are not provided in the summary.
Adaptive explanations reduce over-reliance on the AI system.
Reported result from the prototype experiment indicating lower over-reliance when adaptive explanations are used; specific experiment design and sample size not provided in the abstract.
Experimental results from a medical decision-support prototype reveal that adaptive explanations improve trust accuracy.
Empirical evaluation using a medical decision-support prototype comparing adaptive explanations to a baseline; the paper reports experimental results but does not state sample size or full methodological details in the provided text.
We introduce a framework that tailors explanations to user expertise levels, integrating human factors research with AI interface design.
Paper presents a conceptual/framework contribution describing adaptive explanation tailoring informed by human factors and interface design; this is a proposed framework rather than an empirical result.
The paper outlines evidence-based design principles that balance automation efficiency with human oversight, structural clarity, and adaptive learning.
Prescriptive principles synthesized from management literature and technical practice; labeled 'evidence-based' by the authors but no new empirical evaluation or quantitative validation reported in the paper.
The paper proposes a management-informed vocabulary for agentic systems to improve conceptual clarity and governance.
Proposal of terminology and conceptual vocabulary within the paper; presented as a design contribution rather than empirically validated terminology.
By integrating organizational design principles with technical implementation practices, practitioners can move agentic AI from experimental art toward evidence-based organizational capability.
Normative argument and proposed integration strategy in the paper; no reported controlled studies or measured outcomes demonstrating this transition.
Management theory—spanning boundary objects, spans of control, decision rights allocation, and organizational architecture—offers essential conceptual tools for designing and governing multi-agent systems.
Theoretical synthesis and argument drawing parallels between management concepts and multi-agent system design. The paper proposes the use of these management constructs but does not present empirical evaluation.
Productivity gains from AI are significant at the firm level.
Synthesis of firm-level empirical studies in the SLR reporting positive impacts of AI adoption on firm productivity metrics.
Net employment outcomes from AI adoption are positive overall but unequally distributed across workers/occupations.
Aggregate conclusion drawn from the 78-study SLR indicating more studies report net positive employment effects while highlighting distributional heterogeneity across skill/occupation groups.
AI generates new AI-complementary roles.
Synthesis of studies in the SLR reporting job creation and task-complementarity effects where AI augments worker tasks and creates new roles.
Human-only teams identified significantly more major coding errors than AI-assisted teams.
Reported error-detection outcomes from randomized experiment; summary explicitly states human-only teams identified significantly more major coding errors.
AI-assisted teams (using ChatGPT as a collaborative tool) achieved a 91% reproduction rate.
Reported reproduction outcome from randomized experiment (teams using LLM assistance); summary statement gives 91% reproduction rate for AI-assisted teams.
Human-only teams achieved a 94% reproduction rate when attempting to reproduce published quantitative social science results.
Reported reproduction outcome from randomized experiment (teams attempting to reproduce published results); summary statement gives 94% reproduction rate for human-only teams.
The findings provide city-level evidence supporting coordination of AI development with green computing infrastructure and low-carbon governance.
Policy implication drawn in abstract from the empirical results and heterogeneity/interaction findings.
Heterogeneity analyses reveal stronger AI–green productivity effects in cities with higher green computing capacity, stronger industrial foundations, and weaker resource-environmental constraints.
Reported heterogeneity analyses in abstract splitting sample by city characteristics (green computing capacity, industrial foundation strength, resource-environmental constraint levels).
Mechanism-consistent evidence suggests three transmission channels: green knowledge recombination, intelligent regulation of carbon energy flows, and green value-chain coordination.
Mechanism analysis reported in abstract indicating evidence consistent with three specified channels; likely based on auxiliary regressions or mediation-style tests.
Green computing capacity is associated with a stronger AI–green productivity relationship.
Heterogeneity / interaction analysis reported in abstract showing the AI effect on green productivity is larger in cities with greater green computing capacity.
The positive AI–green productivity result remains robust after alternative measurements, sample restrictions, winsorization, lagged regressors, and Bartik instrumental variable estimations.
Reported robustness checks in abstract including alternative measures, sample restrictions, winsorization, lagging regressors, and use of a Bartik-style instrumental variable.
There is a robust positive association between AI technological development and urban green productivity.
Panel regression analysis on 287 Chinese cities (2005–2023) reported in abstract; association described as robust.
Overcoming threshold constraints and harnessing spatial spillovers are essential to foster coordinated development and realise the full potential of green productivity.
Policy/recommendation drawn from the empirical findings (discussion/conclusion in paper); not itself an empirical estimate—no sample size or direct test reported in summary.
The spatial spillover effect of DIA–DIT synergy on green productivity depends on development stage: in less developed regions the synergy yields positive spatial spillovers for GP.
Spatial Durbin model estimates reported in the paper indicating positive spillover coefficients for less developed regions; summary provides no numeric coefficients or sample size.
When regional knowledge breadth and innovation exceed their respective thresholds, their favourable impact on green productivity is amplified.
Threshold model results reported in the paper showing positive amplification of GP once knowledge breadth and innovation pass identified thresholds; no numeric thresholds or sample sizes in summary.