Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Findings show that productivity gains associated with AI are strongly influenced by organisational readiness, including digital maturity, workforce capabilities, governance quality, and institutional coordination.
Synthesis of results from the systematic review of 68 empirical studies assessing productivity outcomes, methodological quality, effect sizes, and contextual factors.
Technology adoption alone is insufficient for improving SME export performance; sustained innovation efforts and productivity-enhancing routines are the more decisive foundations of export competitiveness.
Interpretation and conclusion drawn from the study's regression results (significant positive effects for innovation and productivity, non-significant effects for digital transformation and AI) as stated in the abstract.
Using industry-level panel data (63 observations) and pooled OLS and fixed-effects estimations, the analysis evaluates internal capability factors and external structural influences for Vietnam's manufacturing SMEs over 2015–2023.
Study design and methods as described in the abstract: industry-level panel, period 2015–2023, 63 observations, pooled OLS and fixed-effects.
The explanatory power of the model is substantial (R² between 0.642 and 0.701), suggesting capability-related factors account for a meaningful share of export variation across industries.
Reported model R² range for the panel regressions (pooled OLS and fixed-effects) in the abstract.
Labor productivity exerts a significant positive effect on export performance (β = 24.57, p < 0.05).
Industry-level panel regression analysis (pooled OLS and fixed effects) on Vietnam manufacturing sector, 2015–2023; reported coefficient and p-value in abstract.
Innovation is positively associated with export performance (β = 45.61, p < 0.01).
Industry-level panel regression analysis (pooled OLS and fixed effects) on Vietnam manufacturing sector, 2015–2023; reported coefficient and p-value in abstract.
Implementing the engineering mechanism 'Baseline-Log Physical Separation' reduced AI Instructions volume by ~75% in the same project.
Reported before/after measurement in the Bang-v3 project after deploying the Baseline-Log Physical Separation mechanism; authors report ~75% reduction.
Treating embodied memory as depreciating capital and pricing that stock with a single endurance shadow price η makes cost-minimizing placement across a RAM / on-board NVM / cloud hierarchy a threshold in a wear-augmented per-byte index.
Analytical/theoretical model developed in the paper showing that introducing a single shadow price η yields a threshold routing policy based on a wear-augmented per-byte index. No empirical sample size reported.
A robot's flash endurance is a non-renewable stock: every persisted write spends one of a few thousand program/erase cycles and never refills, yet no fielded robot memory system prices which memories are worth an erase cycle.
Descriptive/observational assertion in the paper; cites typical flash P/E limits ("a few thousand program/erase cycles") and claims absence of memory-level pricing in deployed robot systems. No sample size reported.
Agentic AI can become a productivity lever when implemented as a human-centered capability with responsibility and accountability retained by people.
Paper's concluding recommendation (argumentative; no empirical evaluation or sample reported).
For small and medium sized companies, agentic systems can improve the use of organizational knowledge.
Paper's conceptual claim about better leveraging organizational knowledge (argumentative; no empirical sample).
For small and medium sized companies, agentic systems can accelerate routine processes.
Paper's argument about process speedups in SMEs (conceptual reasoning; no experimental data reported).
For small and medium sized companies, agentic systems create potential to reduce administrative burden.
Paper's argument about expected benefits for SMEs (conceptual reasoning; no reported empirical sample or trial).
We introduce mojo-deterministic, an open-source library of reproducible reduction kernels.
Paper announces the release/introduction of an open-source library (mojo-deterministic) providing reproducible reduction kernels; likely accompanied by repository link or code artifacts in the full text.
On Apple Silicon, Mojo demonstrates 20x to 180x speedups over pure Python on directly measured kernels.
Reported benchmark results on Apple Silicon comparing Mojo to pure Python on directly measured kernels; exact kernel count and experimental details not provided in the abstract.
Its MLIR compilation infrastructure further allows a single codebase to target scalar, SIMD, multicore, and GPU execution, reducing the translation bottleneck between research and production.
Technical claim in paper about Mojo's MLIR-based compilation pipeline enabling multiple backend targets from one codebase; described as reducing translation work.
While closing the Python-to-C++ performance gap, Mojo uniquely combines native interoperability with the low-level systems control required to construct bit-exact deterministic kernels.
Paper claim, supported by the authors' benchmarks and description of language features (native interop and low-level control); specific benchmark details partly provided elsewhere in the paper.
This article surveys Mojo, Modular's 2026 Python-like systems language, as a structural response for capital markets engineering.
Paper declares itself a survey of the Mojo language applied to capital markets engineering; descriptive statement rather than empirical evidence.
Mechanism analysis indicates AI operates primarily through R&D absorption capacity, agricultural productivity improvements, and land resource optimization rather than through direct volumetric expansion of biofuel inputs.
Mechanism analysis reported in the paper (additional regressions/mediation tests) showing associations between AI measures and R&D absorption, ag productivity, land optimization indicators; direct volumetric channels found weaker.
Venture capital investment in AI technologies generates a complementary but more modest positive effect on biofuel production (cumulative elasticity: 0.076) over a 1‑year horizon.
Panel FGLS regressions with distributed lags showing VC in AI as an explanatory variable; reported cumulative elasticity and timing.
AI-related scientific publication volume exerts a positive and statistically significant effect on biofuel production, with a cumulative elasticity of approximately 0.47 materializing predominantly through a 2‑year lag.
Panel FGLS regressions with distributed lags (controls for policy shocks and time effects); significance reported in main regression results.
Practitioners can adopt oracle-aware quality checks to more accurately evaluate agent-authored contributions.
Recommendation derived from empirical findings (prevalence of weak/no oracles and the link between strong oracles and higher adjusted merge likelihood).
A regression analysis adjusting for agent, PR size, repository popularity, task type, and language shows strong oracles significantly improve merge likelihood (OR = 1.28, p < 0.001).
Multivariate regression (logistic) on the study dataset controlling for listed covariates; reported odds ratio and p-value.
Recent studies report more than 932,000 agent-authored PRs across more than 116,000 repositories.
Cited prior empirical studies (reported counts) as stated in the paper's introduction/related work.
Experiments demonstrate the efficiency of the proposed probing strategy (i.e., it reduces cost/uncertainty efficiently in practice).
Empirical results on synthetic and real-world benchmarks claimed in the paper; specific numeric improvements or sample sizes are not provided in the excerpt.
Extensive experiments on synthetic and real-world benchmarks validate the theoretical predictability regimes described in the phase diagram.
Empirical experiments reported in the paper on both synthetic data and real-world benchmarks; the excerpt does not include dataset names, counts, or sample sizes.
Based on these dynamics, the paper derives a budget-optimal probing principle for pre-hoc performance prediction.
Theoretical derivation of a probing/budget allocation principle (methodological contribution; supported by subsequent experiments according to the text).
Prediction risk decomposes into two components: an intrinsic limit (static data-model compatibility) and a reducible optimization variance.
Theoretical decomposition derived in the paper (analytical derivation/proof; no empirical sample size in excerpt).
Pre-hoc performance prediction can be formally formulated as a stochastic estimation problem under information constraints.
The paper states this as its theoretical formulation/approach (theoretical/mathematical modeling; no empirical sample size).
Pre-hoc performance prediction offers a critical solution to substantially reduce the expense of fine-tuning LLMs.
Paper proposes pre-hoc performance prediction as a method and claims it can substantially reduce fine-tuning costs; later supported by experiments (described as extensive) on synthetic and real-world benchmarks (no numeric reductions given in the excerpt).
The proposed decision-centric portfolio framework provides a pathway to resolving the AI-investment paradox by linking AI investments to identifiable, governable, and accumulative sources of business value.
Synthesis/concluding claim based on the theoretical framework developed in the paper (AIPNs + Expected Net Benefit + staging and portfolio assembly); no empirical test of whether using the framework actually resolves the paradox is provided in this paper.
AIPNs can be staged using real options logic and assembled into a broader portfolio using risk–return principles to guide investment sequencing and allocation.
Conceptual/methodological claim in which the authors show how real-options reasoning and portfolio theory apply to staged investment in AIPNs; presented as framework guidance without empirical implementation in this paper.
Node-level value of an AIPN can be formalized through Expected Net Benefit.
Theoretical formalization presented in the paper (mathematical/analytic definition of Expected Net Benefit at the node level); no empirical estimation reported.
Introducing AI-Investable Process Nodes (AIPNs) — bounded decision points in workflows where AI can alter expected outcomes — enables ex ante assessment of benefits, risks, and costs.
Conceptual framework and definition introduced by the authors; formal description of AIPNs and argumentation showing how they permit ex ante assessment (no empirical validation reported).
At ambiguous tags where a single-pass baseline silently mis-binds 75.0% of the time, FacProcessTwin defers to the operator and mis-binds none.
Comparison between a single-pass automated baseline and FacProcessTwin's human-in-the-loop governance on ambiguous tags in the case study; baseline mis-bind rate reported as 75.0%, FacProcessTwin mis-bind rate reported as 0%. Sample size of ambiguous tags not stated in abstract.
FacProcessTwin builds each twin in roughly a sixth of the manual time (i.e., about 1/6 the time a manual build takes).
Time-to-build comparison versus manual baseline reported in case study covering 16 process flows.
FacProcessTwin generates these process models accurately, achieving a mean F1 of 95.2% against ground truth.
Quantitative evaluation against ground-truth labels in the case study (metric: mean F1). Sample context: 16 production process flows from single manufacturer.
The generated model and its data bindings are rendered as an interactive process diagram through which manufacturing personnel can monitor and correct the system's autonomous decisions, including resolving uncertainty at safety-critical binding steps.
System user-interface and human-in-the-loop governance functionality described in paper (implementation claim).
FacProcessTwin generates a complete process model and automatically binds its process steps to live operational data.
System functionality described in paper (implementation claim).
FacProcessTwin leverages a large language model (LLM) to reduce process twin development time by building a process twin from a plant's process documentation and natural-language input from an operator.
System design and implementation described in paper (methodological/system claim).
Process twins provide real-time representations of entire production processes and have the potential to drive efficiency gains across the whole process.
Conceptual argument presented in paper (definition and motivation of process twins); no empirical test of economy-wide efficiency gains provided in abstract.
The study provides critical theoretical and practical insights for firms integrating AI into high-level governance frameworks.
Claim about the contribution of the paper (theoretical and practical insights); this is a statement of scope/contribution rather than an empirical result—no evidence metrics supplied in the summary.
By fostering collaborative intelligence, organizations can leverage GenAI’s computational reach to improve decision outcomes.
Paper argues as a practical implication that collaborative intelligence enables firms to use GenAI's computational capacity to enhance decision outcomes; no measured effect sizes or sample reported in the summary.
AI's role has shifted from a peripheral tool to a central architect in strategy development.
Framed as an interpretation of the study's findings about role-change in governance; no longitudinal adoption data or counts reported in the summary.
AI can surpass human proficiency in complex domains.
Presented in the paper's findings as an asserted empirical/general conclusion; the summary does not include experimental design, comparative metrics, or sample size.
GenAI agency functions as a mediator between human skill development and algorithmic trust.
Paper explicitly states this mediation relationship as part of its theoretical model; the summary provides no empirical mediation analysis details (no N, no coefficients).
Human-machine shared intentionality enables navigation of organizational complexity.
Framed in the paper as a conceptual mechanism (shared intentionality) that helps organizations manage complexity; summary does not report empirical tests or sample details.
The emergence of "Joint Agency" in corporate governance, where generative AI (GenAI) and human leaders collaborate, enhances Strategic Decision Quality (SDQ).
Paper presents this as a central theoretical claim and summarizes findings supporting it; no empirical sample size, statistical tests, or controlled experiment details provided in the summary.
Post-crisis, output-target pressure can produce a false-correction loop in which agents patch AI failures with more AI.
Model dynamics and a formal proposition in the paper describing post-crisis behavior. No empirical data.
Rational agents incur positive cognitive debt because the costs are deferred, partially external, and masked by short-run productivity gains.
Analytical proof/proposition(s) in the formal model (the paper states this as a derived result). No empirical sample.