Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We successfully deployed GrowthGR on Taobao's production platform, achieving a substantial 5.3% lift in new item GMV.
Reported result from a production deployment / online A/B testing on Taobao (deployment and observed lift claimed in paper); no sample size or experimental details provided in the excerpt.
The Multi-Value-Aware Generative Retrieval (MultiGR) module, built on a semantic-ID-based generative retrieval architecture, leverages structured samples with search cascade signals and adopts a Multi-Value-Aware Policy Optimization (MoPO) training paradigm to align with multi-stage online values while explicitly balancing short-term transactional value and long-term growth potential estimated by ItemLTV.
Methodological description in the paper (design of MultiGR and MoPO); no empirical results cited in this sentence.
The Item Long-term Transaction Value Prediction (ItemLTV) module employs counterfactual inference to quantify the long-term value increment attributable to a single user interaction.
Methodological description in the paper (design of ItemLTV module); no experimental quantification provided in excerpt.
We propose a Multi-Value-Aware retrieval framework (GrowthGR) tailored for e-commerce search, designed to better align with the cascaded online values across different stages of the search system while balancing immediate conversion and long-term item growth.
Methodological contribution described in the paper (system/algorithm proposal); no empirical evaluation details in this sentence.
Root-cause-diagnosis accuracy rises from 75% to 100% when agents have causal grounding (Causely) in the active-fault scenario.
Reported diagnostic accuracy rates from benchmark experiments comparing runs without Causely (75%) and with Causely (100%) on the active-fault scenario.
Causal grounding lowers direct API cost per run by 57%.
Reported percent reduction in direct API cost per run from benchmark experiments with vs. without Causely.
Causely compresses the investigation footprint by 4.8× (in the active-fault scenario).
Reported multiplicative reduction in investigation footprint from benchmark experiments comparing runs with vs. without Causely.
On the active-fault scenario, causal grounding reduces mean tool-call count by 78%.
Benchmark experiment results comparing tool-call counts with vs. without Causely on the active-fault scenario (quantitative percentage reported in paper).
On the active-fault scenario, causal grounding reduces mean token consumption by 60%.
Benchmark experiment results comparing token consumption with vs. without Causely on the active-fault scenario (quantitative percentage reported in paper).
On the active-fault scenario, causal grounding reduces mean time-to-diagnosis by 63%.
Benchmark experiment results comparing agent runs with vs. without Causely on the active-fault scenario (quantitative percentage reported in paper).
Causely transforms raw telemetry into a live, queryable model providing the semantic and causal foundation AI agents require to diagnose, evaluate impact, and act safely in production.
System design / implementation claim in the paper (supported by downstream benchmark evaluation described elsewhere in the paper).
Causely is a causal intelligence layer that maintains a structured representation of environment topology, attribute dependencies, and causal relationships anchored to an ontological representation of the managed environment.
System design / implementation claim in the paper describing the Causely architecture.
Managerially, firms should pair GenAI access with short AIC micro-training and simple standard operating procedures (SOPs) to capture value consistently and avoid uneven adoption outcomes.
Authors' managerial recommendation drawn from experimental findings that AIC predicts gains and that scaffolding reduces variance; recommendation is an interpretation/synthesis rather than a directly tested organizational field intervention.
A scaffolding intervention (conceptual maps) reduced outcome variance, indicating that standardized workflows can mitigate inequality in AI-mediated performance.
Experimental inclusion of a scaffolding intervention (conceptual maps) and reported reduction in variance of outcomes among participants receiving scaffolding in conjunction with GenAI access.
Improvements were not predicted by GPA or prior knowledge, but were predicted by AI Interaction Competence (AIC) — the ability to elicit, filter, and verify model outputs.
Regression/subgroup analyses reported in the experiment linking improvements in task performance to measured predictors (GPA, prior knowledge, AIC); authors report null association for GPA/prior knowledge and positive association for AIC.
On average, GenAI access significantly increased task performance.
Reported randomized controlled experiment comparing task performance between LLM-assisted group and traditional-resources group; authors state the average increase was statistically significant.
The policy architecture required to escape the trap (targeting trust, sequencing, and team-level adoption) is characterised.
Model-derived policy prescriptions identifying interventions (trust-building, sequencing, team-level targeting) necessary to shift equilibria toward genuine adoption; theoretical argumentation. No empirical trial or sample.
Conditions are derived under which sustained but imperfect adoption pressure is welfare-improving.
Analytical derivation within the model framework characterising parameter regions where persistent imperfect adoption increases welfare (model-defined welfare metric). Theoretical analysis; no empirical sample.
A cost ratchet dynamic implies that failed adoption attempts permanently lower barriers even when embedding fails.
Model component introducing a cost-ratcheting mechanism; analytical/simulation results showing permanent barrier reductions following failed attempts. Theoretical model; no empirical sample.
Benchmark comparisons of multiple LLM backends (Granite-Docling, Mistral-Small, DeepSeek-OCR) were performed to provide practical insights for production deployment.
Authors state they performed benchmark comparisons of multiple LLM backends (listed in abstract); specifics of metrics and sample sizes not given in abstract.
A comprehensive sustainability analysis shows that the hybrid AI+HITL approach reduces CO2 emissions by 69%, energy consumption by 69%, and water usage by 63% compared to traditional manual processing.
Authors report a sustainability analysis comparing hybrid AI+HITL approach to traditional manual processing (details not provided in abstract).
Prompt Fine Tuning with Feedback Inheritance (PFTFI) is a novel approach introduced in this work.
Authors explicitly introduce PFTFI as part of their approach (stated in abstract).
The system integrates five specialized agents—Classificator, Splitter, Parser, Extraction, and Validator—together with a Human-in-the-Loop mechanism and a Prompt Fine Tuning with Feedback Inheritance (PFTFI) approach.
Authors' architectural description in the abstract specifying the five agents, HITL mechanism, and PFTFI approach.
MADP combines deep learning-based classification and parsing with large language model extraction while maintaining accuracy through selective human validation.
System description in paper asserting integration of DL classification/parsing, LLM extraction, and selective human validation; supported by system evaluations reported elsewhere in abstract.
Ablation evaluation on a stratified 100-document subset demonstrates that the full MADP configuration with Human-in-the-Loop supervision attains 98.5% document-level accuracy.
Ablation evaluation on a stratified subset of 100 documents (5 documents per each of 20 supplier/document-type categories) reported by authors.
Only 3% of documents required non-AI fallback in the production deployment.
Same production deployment on 955 documents (stated in abstract).
Production deployment on 955 real-world documents processed through January 2026 achieves a 97.0% full-pipeline automation rate.
Reported production deployment on 955 real-world documents processed through January 2026 (stated in abstract).
Operational analysis on a production use-case scenario of 100,000 invoices per year indicates a potential reduction of Full-Time Equivalent (FTE) requirements by approximately 70%.
Operational analysis reported by authors on a production use-case scenario involving 100,000 invoices per year (stated in abstract).
A year-long pilot across three clinical sites executed 8,728 cohort-enrolled workflow runs with a 97.08% completion rate under an early prototype without the verified-core subsystem.
Reported evaluation: year-long pilot conducted across three clinical sites, total workflow runs = 8,728, reported completion rate = 97.08%; prototype lacked verified-core subsystem.
Swimlanes make trust boundaries explicit, separating verified logic from external systems, human judgment, and AI decisions.
Design description in the paper explaining swimlane use to delineate trust boundaries between system components and humans/external systems/AI.
At runtime a durable engine records outcomes in an append-only event log and can enforce contracts at system boundaries, supporting replay, retries, and audit.
System architecture description in the paper describing runtime engine features (append-only log, enforcement, replay/retry, audit support).
At compile time GraphFlow restricts diagrams to produce reusable automations whose contracts (preconditions, postconditions, and composition obligations) are intended to be proof-checked before admission to a shared library.
Design/specification claim in the paper describing compile-time restrictions and proof-checked admission model (implementation/design detail).
GraphFlow treats workflow diagrams as the executable specification — a single artifact defining data scope, execution semantics, and monitoring — to address the gap between durable execution and semantic correctness.
System design description in the paper explaining GraphFlow's design philosophy and intended role of diagrams as executable specifications.
Existing workflow platforms provide durable execution and observability.
Author statement in background/motivation describing properties of existing workflow platforms.
Context engineering (programmatic state abstraction and clean task decomposition) is generally more cost-effective than deeper per-agent deliberation.
Cost-effectiveness measured as returns per token spent (RPTS) across configurations that vary context representation and deliberation; results from the 3,475-episode controlled study indicate context changes yielded larger returns per token than adding deliberation tools.
Programmatic state abstraction delivers the largest returns per token spent (RPTS), improving mean return by up to 76% over raw observations.
Controlled empirical study in the CybORG CAGE-2 POMDP environment comparing context representations (raw observations vs. deterministic state-tracking layer with compressed history) across five model families, six models, and twelve configurations with token-level cost accounting (3,475 episodes).
For AI datacenter design, the relevant planning objective is not installed megawatts, but deployable capacity over time.
Conclusion/recommendation drawn from the paper's modeling results and analysis (argument that installed MW is a poor planning metric compared to time-varying deployable capacity).
The framework combines projection models for GPU, compute, and storage deployments with operational factors grounded in production data from Microsoft Azure.
Method claim: framework integrates projection models and operational data from Microsoft Azure (production data grounding); stated in the paper's methods summary.
We develop a framework for evaluating datacenter power delivery designs using throughput, power, and cost metrics over realistic arrival, oversubscription, and decommissioning sequences.
Methodological claim describing the paper's core contribution: a simulation/evaluation framework combining throughput, power, and cost metrics with arrival/oversubscription/decommissioning sequences; based on the authors' implementation (details and data referenced in the paper).
Designs must remain efficient over long datacenter lifetimes and multiple hardware generations.
Normative/design recommendation motivated by long asset lifetimes and evolving hardware density; stated as a requirement in the paper.
Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027.
Projection models for GPU deployments described in the paper (projection models combined with industry deployment assumptions); specific provenance referenced in the abstract but no sample size reported.
The study's findings offer actionable insights for managers and policymakers to leverage AI for sustainable organizational growth while safeguarding employee well-being.
Authors' concluding statement based on survey findings and analytical results.
Successful human–AI collaboration requires a human-centric approach that balances technological advancement with workforce development, ethical governance, and organizational support.
Study conclusion/recommendation based on survey findings (perceptions of opportunities and challenges) and analytical results (correlation/regression).
Human–AI collaboration reduces employees' routine workload.
Respondent perceptions collected via the structured questionnaire and analyzed with descriptive statistics and regression in SPSS.
AI-based systems support better decision-making by providing data-driven insights, allowing employees to focus on higher-level cognitive and strategic activities.
Survey responses (structured questionnaire) analyzed with SPSS (correlation and regression analyses) reporting perceived support for decision-making.
Human–AI collaboration significantly enhances workplace efficiency and productivity by reducing routine workload and improving accuracy and speed in task execution.
Primary data from employees in AI-enabled organizations collected via a structured questionnaire (5-point Likert); analyzed with SPSS using descriptive statistics and regression analysis.
Continuous, simulation-driven prompt optimization is both tractable and necessary for reliable enterprise conversational AI at scale.
Concluding claim in abstract: 'Our results suggest that continuous, simulation-driven prompt optimization is both tractable and necessary...'.
PRISM is designed to run on a scheduled basis (daily), treating LLM behavioral drift as a first-class reliability concern.
Design statement in abstract describing scheduled daily runs to monitor behavioral drift.
PRISM diagnoses root causes of failures and surgically repairs the prompt, iterating until all tests pass.
Methodological description in abstract stating diagnosis and iterative repair loop until tests pass.
PRISM simulates full multi-turn conversations against a platform-faithful LLM environment and evaluates pass/fail using an LLM-as-judge.
Method/architecture claim in abstract describing simulation of multi-turn conversations and LLM-based judging.