Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
DOMUS is replicable digital public infrastructure: a modular, cloud-native Software-as-a-Service architecture that can be deployed across other UK boroughs and adapted to other public administration tasks characterised by scarcity, rule-bound eligibility, and high stakes.
Design and architectural claim in the paper about modular cloud-native SaaS architecture and intended replicability; this is an assertion about generalisability rather than evidence from multi-site deployments.
The deployment maintained statutory compliance and role-based accountability.
Paper asserts that DOMUS operations preserved statutory compliance and role-based accountability during the pilot (claimed based on system design and pilot evaluation; no specific compliance audits or metrics provided in the provided text).
Results indicate high staff satisfaction.
User feedback / staff satisfaction reported from the pilot deployment (paper states high satisfaction but does not report sample size or survey statistics in the provided text).
Results indicate improved adherence to key placement constraints.
Pilot evaluation results reporting better adherence to placement constraints (e.g., bedroom need, affordability, accessibility) under DOMUS compared to manual workflows (no quantitative metrics provided in the provided text).
Results indicate substantial reductions in search time.
Findings from the pilot deployment comparing DOMUS-assisted search time to manual search workflows (paper reports reductions but does not provide numeric effect size in the provided text).
Household and property attributes are encoded into policy-consistent representations prior to AI-assisted ranking and explanation.
Technical design detail in the paper describing preprocessing/encoding pipeline used by DOMUS before AI ranking and explanation.
The system combines transparent, rule-based filtering with large language model-assisted search to standardise the application of bedroom need, affordability thresholds, geographic preferences, and accessibility requirements, while preserving officer discretion and audibility.
Design and functionality description; statement that DOMUS uses rule-based filtering plus LLM-assisted search to standardise applications of policy rules and preserve discretion/audibility.
DOMUS integrates household case records, policy-constrained affordability and suitability rules, and live private-rental listings within a single governance-aligned workflow.
System architecture and functionality described in the paper; design claim about integrated data sources and workflow alignment.
The paper documents the creation and use of DOMUS, a cloud-based, AI-enabled decision-support system built from scratch at the University of East London and customised for the needs of London Borough of Newham to support statutory Temporary accommodation placement.
Implementation and deployment description in the paper (system development and customised pilot deployment for Newham).
LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.
Authors' claim about the benchmark's intended properties (reproducibility, low cost) based on its web-based simulated design and browser operation; no cost comparison or reproducibility study reported in the excerpt.
LabOSBench is a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators that operates directly via a browser and thus avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation.
System and benchmark description in paper; architectural/design claim (web-based simulators, browser operation, avoidance of OS virtualization, supports flexible configuration and execution-based evaluation).
The authors release their code and agent trajectories to support future research.
Statement in the paper indicating that code and recorded agent trajectories are publicly released.
Most evaluated models achieve positive net income in the CoffeeBench simulation.
Reported experimental results stating that the majority of evaluated LLMs produced positive cumulative net income.
Across several recent open-weight and proprietary LLMs, all models outperform a passive baseline that takes no actions.
Empirical evaluation comparing multiple recent open-weight and proprietary LLMs to a passive baseline (passive baseline defined as taking no actions); paper reports that every tested model outperformed that baseline.
We introduce CoffeeBench, a benchmark for evaluating LLM agents in a long-horizon multi-agent economy composed of heterogeneous firms.
Design and release described in the paper: authors state they developed CoffeeBench as a benchmark for long-horizon multi-agent economic evaluation.
An agnostic model is formed (combining theoretical constructs, empirical evidence, and practical applications) enabling organizations to have the accountability, operational and human oversight needed to embrace responsible AI-enabled automation of enterprise systems and processes.
Paper proposes an agnostic model based on synthesis of theory, empirical evidence, and practice; the excerpt describes the model conceptually without presenting evaluation metrics or sample-based validation.
Artificial intelligence integration requires established governance frameworks, human-in-the-loop verification, and explainable artificial intelligence to ensure compliance with the organization's values and legislation.
Prescriptive recommendation in the paper combining theoretical constructs and practical application guidance; no empirical validation provided in the excerpt.
Systematic AI integration can produce meaningful productivity gains across engineering design, content generation, multimedia creation, and scientific experimentation.
Paper combines theoretical, empirical, and practical examples to claim productivity gains; excerpt does not provide study methodology or sample sizes.
Systematic AI integration can produce meaningful improvements in configuration accuracy.
Asserted by paper based on examples from engineering, content, multimedia, and scientific workflows; excerpt contains no measurement details or sample sizes.
Modern engineering design, content generation, multimedia creation, and scientific experimentation conducted by organizations show that meaningful savings in development time can be realized by the systematic integration of artificial intelligence technologies into existing quality management systems.
Paper synthesizes theoretical constructs, empirical evidence, and practical applications to assert time-savings across multiple domains; no specific study design, sample size, or quantified effect provided in the excerpt.
Measurable system reliability improvements have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper reports improvements in system reliability associated with LLM integration; excerpt lacks details on measurement approach, sample, or magnitude.
Measurable defect improvements (reduction in defects) have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper claims empirical reductions in defects linked to LLM integration; no methodological details or sample size provided in the excerpt.
Measurable productivity improvements have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper asserts empirical/measurable productivity improvements attributable to LLMs integrated with enterprise platforms; the excerpt provides no details on study design, measurement method, or sample size.
Organizations should consider governance and quality management when introducing generative AI.
Normative recommendation by the paper, presented as best-practice guidance; based on theoretical constructs and practical applications described, not an empirical test in the excerpt.
Automated orchestration, code generation, and generative creativity have been introduced as well.
Descriptive claim in paper indicating the introduction/adoption of specific generative-AI capabilities; no empirical details or sample size given in the excerpt.
Generative AI systems have been incorporated into innovation and process optimization in organizations.
Stated in paper as an observed trend / descriptive claim; no specific study design, sample size, or empirical method reported in the excerpt.
The authors release an updated version of the WorkBench benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024.
Statement of release in the paper (announcement of updated benchmark contents); verifiable by checking the authors' repository or supplementary materials referenced in the paper.
The rise of open-weight models has drastically lowered costs for a performance level that was previously only accessible to proprietary models.
Comparative cost analysis reported by authors linking open-weight model availability to lower costs for a given performance level; the excerpt does not include specific cost figures, sample sizes, or methodology.
Several classes of error have been totally eliminated from frontier agents on WorkBench.
Authors' qualitative analysis of error types observed in WorkBench evaluations between 2024 and 2026; the excerpt does not list which error classes, how elimination was measured, or sample sizes.
On WorkBench, capability and safety go together rather than trade off: models that finish the most tasks also do the least unintended damage.
Comparative analysis of agent performance on the WorkBench benchmark showing association between higher task completion rates and lower rates of unintended harmful actions; specific statistical measures and sample size not provided in the excerpt.
In June 2026 the best agent to date, Claude Opus 4.8, completes 89% of tasks on WorkBench.
Reported evaluation result on the WorkBench benchmark (June 2026) comparing agent task completion rates; exact sample size not stated in the excerpt.
In March 2024 the best agent on WorkBench, GPT-4, completed 43% of tasks.
Reported evaluation result on the WorkBench benchmark (March 2024) comparing agent task completion rates; exact sample size not stated in the excerpt.
Moving beyond experimental phases requires high-impact use cases and decentralized governance.
Paper emphasizes (argues) that scaling AI past experiments depends on choosing high-impact use cases and adopting decentralized governance; presented as recommendations rather than empirically validated findings in the summary.
The research offers guidance for bridging the gap between technical success and business impact through operational mitigation strategies.
Paper provides proposed operational strategies and guidance (prescriptive content); no evidence of empirical testing given in the summary.
The verifier does buy sound coverage: covering all unseen valuable statements while asserting only valid ones is possible with it, impossible without it; it relocates unavoidable errors from false to trivial.
Formal proofs within the paper's model showing existence (with verifier) and impossibility (without verifier) results; analysis of error types shifting from false statements to trivial (verifiable-but-not-valuable) statements.
A data-driven culture supports both AI deployment depth and breadth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (data-driven culture measured and shown to be associated with both depth and breadth).
Digital infrastructure drives AI deployment breadth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (digital infrastructure included as predictor of the breadth dimension).
AI-skilled human capital drives AI deployment depth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (human capital measures included as predictors of the depth dimension).
Firm performance is positively associated with the interaction between AI deployment depth and AI deployment breadth (consistent with a complementarity logic).
Empirical analysis using archival microdata from 770 large Spanish firms; staged OLS regression models testing the interaction between measured AI deployment depth and breadth on firm performance.
The contribution is a falsifiable program centered on minimum functional description length and verified-change cost.
Authors' stated research contribution and framing in the paper.
Preliminary QLoRA experiments on Qwen2.5-Coder-14B show that 64,088 canonical trajectories are learnable and suppress tested forbidden-language markers.
Reported preliminary experiment using QLoRA on Qwen2.5-Coder-14B with 64,088 canonical trajectories; empirical claim within the paper.
For supported routine-product distributions, this approach gives a defensible planning target near 100-fold all-in cost reduction (amortized cost per verified correct change), though this is a hypothesis and not a guarantee for all software.
Stated projected limit/planning target in the paper; authors explicitly note these reduction bands are hypotheses and not measured frontier results.
Quotienting software by behavior equivalence under a declared oracle can collapse equivalent encodings into governed representatives with explicit evidence and proof obligations (core hypothesis).
Central theoretical hypothesis / conceptual claim in the paper (no empirical test reported).
We propose an agent-first canonical code, a proof-carrying substrate that rewrites routine product software into canonical behavior profiles, typed change algebra, proof lanes, constrained edit grammars, semantic patch cells, runtime negative memory, and proof-carrying change objects.
Design/proposal presented by the authors (architectural proposal, not an evaluated system).
Human code repositories contain valuable signals such as tests, incidents, migrations, edge cases, product judgment, and operational history.
Asserted observation in the paper describing repository contents and their value; no quantified empirical study provided.
Emotional AI systems may improve organisational productivity when implemented within ethically grounded and transparent frameworks.
Authors' synthesis from the systematic review of the literature (claimed association based on reviewed studies); no quantitative pooled effect size reported in the abstract.
Emotional AI systems may enhance employee engagement when implemented within ethically grounded and transparent frameworks.
Synthesis claim from the systematic review (authors conclude this based on aggregated findings across the surveyed literature); specific studies and counts not reported in the abstract.
Data exhibits spillovers such that data generated by one task can augment the productivity of another task.
Model assumption and formalization of cross-task data spillovers in the analytical framework (theoretical derivation and model structure).
Brick outperforms all tested routers on the evaluated benchmark.
Empirical comparison on the 5,504-query benchmark where Brick's max-quality accuracy (76.98%) is reported as higher than other routers.
Median latency drops from 51.2s to 22.8s when using Brick.
Empirical measurement on the benchmark reporting median latency before and after applying Brick.