Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Exploratory innovation's association with long-term competitive performance operates indirectly through GenAI adoption (mediation).
Survey of 104 Portuguese B2B managers and PLS-SEM showing a mediated pathway from exploratory innovation to performance via GenAI adoption in the estimated model.
GenAI adoption is positively associated with long-term competitive performance.
Survey data from 104 Portuguese B2B managers; association estimated via PLS-SEM in the study's structural model.
Ethical governance is the strongest organisational correlate of long-term competitive performance.
Survey data from 104 Portuguese B2B managers; analysed using Partial Least Squares Structural Equation Modelling (PLS-SEM); reported as a comparative strength of model paths.
By extending traditional technology acceptance models (TAM) with AI-specific dimensions—namely transparency, data quality, and trust—this study contributes to the literature on decision-making in complex systems and offers practical insights for organizations seeking to improve decision effectiveness through AI-based support.
Authors' stated contribution in abstract/introduction; conceptual model extension and empirical tests reported in the paper (survey N = 324 and PLS-SEM results).
Intention to adopt AI-DSS demonstrates a strong association with decision-making efficiency (β = 0.544, p < 0.001).
PLS-SEM path coefficient reported in results (β = 0.544, p < 0.001) linking intention to adopt and decision-making efficiency, estimated from survey data (N = 324).
Perceived usefulness (β = 0.352, p < 0.001), trust (β = 0.311, p < 0.001), and perceived ease of use (β = 0.135, p < 0.05) exert significant positive effects on the intention to adopt AI-DSS.
PLS-SEM path coefficients and significance levels reported for predictors of intention to adopt, based on the questionnaire sample (N = 324).
Perceived ease of use significantly affects perceived usefulness (β = 0.597, p < 0.001).
PLS-SEM estimate reported in paper (β = 0.597, p < 0.001) from the survey of 324 respondents.
Trust positively influences perceived ease of use of AI-DSS (β = 0.482, p < 0.001).
PLS-SEM path coefficient reported in results (β = 0.482, p < 0.001) based on the questionnaire sample (N = 324).
Trust positively influences perceived usefulness of AI-DSS (β = 0.229, p < 0.01).
PLS-SEM path coefficient reported in results (β = 0.229, p < 0.01) from the survey data (N = 324).
Data transparency and quality strongly enhance trust in AI-based decision support systems (AI-DSS) (β = 0.784, p < 0.001).
PLS-SEM estimate reported in results (standardized path coefficient β = 0.784, p < 0.001) based on the survey of 324 respondents.
Evidence-based frameworks for structural redesign that prioritize network density, decision proximity to information sources, and cross-boundary coordination mechanisms are foundational prerequisites for organizational agility.
Concluding synthesis of reviewed literature and empirical cases leading to proposed frameworks. The provided text labels the frameworks 'evidence-based' but does not present quantitative validation or implementation trial results in the excerpt.
The article draws on empirical cases from manufacturing, technology platforms, and healthcare delivery across North America, Europe, and East Asia to support its arguments.
Statement in the article that empirical cases from those sectors and regions were analyzed. The provided text does not specify the number of cases, selection criteria, or methodologies for the case analyses.
Structural reconfiguration enables adaptive behaviors that resist cultivation under traditional pyramid architectures, regardless of cultural interventions.
Claim derived from comparative analysis and empirical case studies referenced in the article; presented as an observation across cases from multiple industries and regions. No explicit statistical tests or counts reported in the provided text.
Flattening hierarchies and redistributing authority to operational edges fundamentally rewires information flow, decision velocity, and collaborative patterns.
Argument based on synthesis of research on organizational modularity and structural determinants of behavior; described as supported by empirical cases across sectors (manufacturing, technology platforms, healthcare). No numerical sample sizes or formal experimental details provided.
Formal structure—specifically hierarchical configuration and decision-making architecture—exerts greater influence on employee behavior than culture change initiatives or compensation redesign.
Synthesis of organizational behavior, network science, and comparative institutional research cited in the article; stated comparison between structural determinants and culture/incentive interventions. No sample size or statistical details reported in the text provided.
Security evaluation across 135 test cases demonstrates 87.5% accuracy on static code safety analysis with zero false positives.
Security evaluation reported in paper across 135 test cases with reported accuracy and false positive rate.
Security evaluation across 135 test cases demonstrates 96.7% accuracy on prompt injection detection.
Security evaluation reported in paper across 135 test cases with reported accuracy metric.
On document intelligence (DocILE), Code Factory achieves the highest line item recognition accuracy (LIR: 80.4%).
Empirical evaluation reported on DocILE dataset of 5,680 invoices; LIR metric reported at 80.4% and described as the highest among compared variants.
Compiled AI reduces token consumption by 57x at 1,000 transactions.
Empirical token-consumption comparison reported in paper (scaling example at 1,000 transactions).
Compiled AI breaks even with runtime inference at approximately 17 transactions.
Cost/efficiency comparison reported in evaluation (function-calling context); break-even point stated in paper.
On function-calling, compiled AI achieves 96% task completion with zero execution tokens.
Empirical evaluation on the BFCL function-calling tasks (reported n=400).
We introduce a system architecture for constrained LLM-based code generation, a four-stage generation-and-validation pipeline that converts probabilistic model output into production-ready code artifacts, and an evaluation framework measuring operational metrics including token amortization, determinism, reliability, security, and cost.
Paper states these three contributions as part of the authors' work (descriptive claim about methods and artifacts presented).
By constraining generation to narrow business-logic functions embedded in validated templates, compiled AI trades runtime flexibility for predictability, auditability, cost efficiency, and reduced security exposure.
Conceptual/systems claim made in paper describing design trade-offs of the compiled AI paradigm (no single empirical test cited in the excerpt).
Experimental evidence confirms that AI tools raise worker productivity.
Statement in paper referencing experimental studies (no specific study, method, or sample size reported in the excerpt).
A lightweight interception layer captures and blocks only the final submission request, ensuring safe evaluation without real-world side effects.
Paper describes an interception layer in the evaluation infrastructure that prevents actual final submissions on production sites.
Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and challenges of real-world web interaction.
Methodological description in the paper: evaluation occurs on live (production) websites rather than offline static sandboxes; supported by reported coverage of 144 live platforms.
The tasks in ClawBench require demanding capabilities beyond existing benchmarks, such as extracting relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and completing write-heavy operations like filling many detailed forms correctly.
Paper description of task types and the capabilities they require; based on the design and composition of the 153 tasks.
ClawBench spans 144 live platforms across 15 categories.
Paper explicitly reports coverage across 144 production websites and 15 task categories (dataset description).
ClawBench is an evaluation framework of 153 simple tasks that people need to accomplish regularly in their lives and work.
Paper states the benchmark comprises 153 tasks (dataset description).
The paper argues for a fundamental decoupling of semantic intent from human-readable representation.
Conceptual/design claim made by the authors as a recommended shift in representation strategy for agentic consumers; presented as argumentation rather than empirically tested in abstract.
We extend the semantic density principle to propose rehabilitation of classical anti-patterns and introduce the program skeleton concept for agentic code navigation.
Design/position claims and proposed constructs presented in the paper (program skeleton concept and re-evaluation of anti-patterns) without empirical validation reported in abstract.
Aggressive compression reduced input tokens by 17%.
Reported numeric result from the controlled experiment comparing compressed logs to other conditions; sample size not specified in abstract.
We propose a key design principle: semantic density optimization, eliminating tokens that carry zero information while preserving tokens that carry high semantic value.
Proposal/design principle presented in the paper; theoretical justification provided and (per paper) subsequently validated by experiment.
These empirical findings provide reference for global governments to optimise artificial intelligence policies for low-carbon urban development.
Paper conclusion interpreting results as policy-relevant and generalisable lessons for governments; based on observed positive association between NAIDPZ and urban GEE.
The impact of the NAIDPZ policy on urban GEE is positively moderated by government attention and public environmental attention.
Reported moderation analysis showing interaction effects between the treatment indicator and measures of government attention and public environmental attention within the DiD framework.
The composite NAIDPZ policy effect increases GEE mainly through promoting green technological innovation and optimising industrial structure.
Mechanism analysis reported in the paper (channel/mediation tests) showing that indicators of green technological innovation and industrial structure optimisation account for much of the policy effect on GEE.
The policy effect on GEE is stronger in inland cities, central-region cities, and non-resource-based cities.
Reported heterogeneity/subgroup analysis within the staggered DiD framework comparing effects across geographic regions (inland vs. others, central vs. others) and city types (non-resource-based vs. resource-based) in the 267-city sample.
The NAIDPZ policy significantly improves urban green economic efficiency (GEE).
Estimated treatment effect from staggered DiD on the 267-city panel (2007–2023) with reported statistical significance and multiple robustness checks mentioned.
ImplicitMemBench reframes evaluation from 'what agents recall' to 'what they automatically enact'.
Paper framing statement positioning the benchmark's conceptual contribution as shifting evaluation focus to implicit, automatic behavior rather than explicit recall.
Top performers were DeepSeek-R1 (65.3%), Qwen3-32B (64.1%), and GPT-5 (63.0%).
Paper lists top model names with reported overall percentage scores from the benchmark evaluation.
The benchmark's 300-item suite employs a unified Learning/Priming-Interfere-Test protocol with first-attempt scoring.
Paper states the suite size (300 items) and describes a unified Learning/Priming-Interfere-Test protocol and that scoring is done on first attempts.
ImplicitMemBench operationalizes three cognitively grounded constructs from cognitive science: Procedural Memory (one-shot skill acquisition after interference), Priming (theme-driven bias via paired experimental/control instances), and Classical Conditioning (CS--US associations shaping first decisions).
Paper description of benchmark design explicitly listing the three constructs and brief operational definitions for each.
We introduce ImplicitMemBench, the first systematic benchmark evaluating implicit memory through three cognitively grounded constructs.
Paper claim of introducing a new benchmark named ImplicitMemBench; it states novelty ('first systematic benchmark') and describes design around three constructs (Procedural Memory, Priming, Classical Conditioning).
Tiny sharing incentives improve models with weak cooperation.
Experimental intervention reported in the paper: adding small sharing incentives and observing improved cooperation among weakly-cooperative models (stated in abstract; no quantitative effect size or sample size provided there).
Explicit protocols double performance for low-competence models.
Experimental intervention reported in the paper: introducing explicit protocols in the multi-agent setup and observing a doubling of performance for low-competence models (stated in abstract; no sample size reported there).
OpenAI o3-mini reaches 50% of optimal collective performance.
Experimental measurement of collective performance for OpenAI o3-mini in the paper's multi-agent setup (value reported in abstract; no sample size provided there).
For non‑tech firms seeking to enhance operational efficiency through digitalization, optimizing internal power structures in response to technological shifts can improve firm performance.
Policy/managerial recommendation based on the study's empirical findings linking digitalization, decentralization, and productivity using China's listed firms data (2009–2020).
Digital technologies operate as an external contingency for non‑tech firms, requiring structural decentralization to align organizational structure with technological shifts.
Theoretical proposition and interpretation of empirical findings in the paper; framed as a contribution to organizational structure theory rather than a separate causal test.
Many non‑technology firms' existing organizational structures fail to accommodate data‑driven digital technologies, creating a need for strategic adaptation to integrate these technologies into business operations.
Argument and literature synthesis presented in the paper motivating the study; descriptive characterization rather than a directly tested empirical claim in the reported analyses.
Shifting power allocation (decentralization to subsidiaries) driven by digitalization significantly enhances firm productivity.
Further (post‑hoc / additional) analyses reported in the paper linking measured shifts in internal power allocation to improvements in firm productivity using the sample of China's listed companies (2009–2020).