Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Knowledge breadth amplifies the positive effect of DIA–DIT synergy on green productivity.
Moderated mediation analysis reported in the paper finds an interaction/moderation by regional knowledge breadth; methods include moderated mediation models; summary gives no sample size or numeric interaction coefficient.
The DIA–DIT synergy promotes green productivity (GP).
Empirical estimation using simultaneous equation models and moderated mediation models linking DIA–DIT interaction to GP (paper summary reports a positive effect); no sample size or numeric effect reported in summary.
The DIA–DIT synergy reduces innovation resource misallocation.
Estimated using the paper's simultaneous equation / mediation framework linking DIA–DIT interaction to measures of innovation resource allocation (methods reported include moderated mediation models); no numeric sample size provided in the summary.
Digital industry agglomeration (DIA) and digital–intelligent technology application (DIT) reinforce each other by combining external resource provision with internal resource orchestration.
Simultaneous equation models and conceptual analysis combining external (regional innovation resources) and internal (enterprise innovation allocation) perspectives, as reported in the paper; no sample size provided in summary.
EvalLoop is packaged as reusable artifacts (playbook, agent specification, template repository) for adoption by other teams.
Paper states that EvalLoop is distributed with supporting artifacts (playbook, agent spec, template repo).
A one-time blind human gate on a finalist panel (4 models, 16 cases) confirms dimensional rankings while resolving multi-criteria deployment trade-offs—yielding a 94% reduction in review burden compared to evaluating the full design.
Reported experimental human-gate evaluation in the paper: 4 finalist models, 16 cases, and quantified review-burden reduction of 94%.
Improvement was concentrated in diagnosed dimensions: Synthesis Power +26.4 percentage points.
Reported per-dimension improvement in the case study after the targeted prompt fix (Synthesis Power +26.4pp).
Improvement was concentrated in diagnosed dimensions: Content Accuracy +16.8 percentage points.
Reported per-dimension improvement in the case study after the targeted prompt fix (Content Accuracy +16.8pp).
A targeted prompt fix improved the best model from 82.6% to 94.6% overall.
Reported performance before-and-after a targeted prompt change in the case study; exact percentages provided (82.6% -> 94.6%).
We validate EvalLoop through a case study on sales intelligence briefing generation (10 models, 3 providers, 18 metrics, 5 dimensions, 3 iterations).
Reported case study design in paper listing 10 models, 3 providers, 18 metrics, 5 dimensions, and 3 iterations.
EvalLoop organizes evaluation around three mechanisms: (1) dimensional metric grouping that decomposes quality into business-relevant dimensions; (2) failure mode classification that categorizes why outputs fail within weak dimensions; and (3) a structured iteration workflow where each evaluation run varies one system variable and compares dimensional profiles before and after.
Methodological description and proposed framework presented in the paper.
The paper offers practical guidance for organizations seeking to improve productivity, decision quality, operational efficiency, and long-term business performance through effective human–AI collaboration.
Paper's discussion and prescriptive recommendations based on its literature synthesis and model; no quantified evaluation of implemented guidance reported in the summary.
The research study adds new knowledge to hybrid intelligence theory and provides organizations with a complete system to implement AI technologies.
Authors' stated contributions in the paper (theoretical addition and prescriptive/system design guidance); presented as the paper's outputs rather than quantified empirical results.
Thriving human–AI collaboration requires, apart from superior AI features, good organizational procedures, reliance on employees, and openness (explainability) of AI systems.
Paper's synthesis and discussion/recommendations based on review of empirical and theoretical studies; presented as design/implementation factors for successful hybrid intelligence.
The researchers developed a task-based adaptive collaboration model which includes hypotheses about how trust, explainability, and task difficulty affect performance results.
Explicit development of a theoretical/model contribution stated in the paper (task-based adaptive collaboration model with hypotheses); described as part of the paper's methods/contributions.
Hybrid intelligence systems make better decisions than separate human or AI systems.
Summary claim in the paper based on synthesis of empirical and theoretical studies from 2021–2026; no single study or sample size provided in the summary.
Hybrid intelligence systems produce up to 60% more work than separate human or AI systems.
Aggregated quantitative finding reported in the paper based on the authors' analysis/synthesis of recent empirical studies (2021–2026); exact constituent studies and sample sizes not specified in the summary.
Organizations now use human–AI collaboration as their primary method to enhance workflow efficiency through hybrid intelligence systems, which replace traditional automation systems.
Statement in paper based on a synthesis of recent empirical and theoretical studies from information systems, organizational behavior, and AI literature (analysis of studies conducted between 2021 and 2026).
Preferred results, based on patents data and first-differenced GMM, suggest that AI adoption already contributes to short-run growth and leads to long-run improvements in standards of living.
Authors' preferred empirical specification (patents-based AI measure) and robustness approach (first-differenced GMM) applied to panel of 35 OECD countries (1995–2017).
The ARDL framework captures the gradual adjustment process and allows incorporation of human capital by interacting it with AI to assess whether AI benefits differ across skill levels.
Methodological claim in paper describing the advantages of the chosen ARDL specification and the use of interactions with human capital.
This study uses a panel ARDL model for 35 OECD countries from 1995 to 2017 to estimate short- and long-run effects of AI adoption on growth and living standards.
Methodological description in the abstract: panel ARDL applied to a sample of 35 OECD countries over 1995–2017.
Technological progress is a key driver of long-term growth and increases in standards of living across generations.
Statement in paper's introduction/background summarizing established literature (no specific new empirical test reported in this abstract).
Where AI likely requires human collaboration, employment rises 4%.
Heterogeneous DiD estimates by exposure type reported in paper; for occupations/industries classified as requiring human collaboration with AI, the estimated employment effect is a 4% increase.
Effects emerge in 2021 when enterprise AI tools entered the market.
Temporal pattern in DiD estimates reported in paper showing treatment effects appearing in 2021, coinciding with the market entry of enterprise AI tools.
A one standard deviation increase in exposure raises output by 7%.
Difference-in-differences estimates using administrative data; exposure measured in standard deviations; reported coefficient = 7% increase in output per one standard deviation increase in AI exposure.
The memory is shared across users: workspaces encapsulating selective memory can be transferred across users with role-based access control, enabling collaborative reuse without redundant specification.
System design and deployment description in the paper indicating implemented role-based workspace sharing in the collaborative workspace platform.
The shared selective persistent memory architecture identifies and retains four categories of reusable context: task specifications, data schemas, tool configurations, and output constraints, while discarding session-specific reasoning traces.
Design and implementation description of the proposed architecture in the paper; implemented in the reported platform.
A replication on four public datasets confirms generalizability, with zero-token refresh succeeding in 12/12 trials.
Replication experiments reported on four public datasets with a total of 12 trials (12/12 successes) as stated in the paper.
Summary-driven generation cuts per-invocation token cost by 97x versus raw data injection.
Reported comparison in the paper between summary-driven generation and raw data injection approaches; exact measurement procedure and sample counts not included in the excerpt.
Zero-token refresh eliminates LLM re-invocation for recurring updates, yielding a 14x task-time reduction.
Reported measurements from the deployed platform experiments in the paper (enterprise scenarios); no per-trial counts provided in the excerpt.
Shared selective persistent memory achieves 96% task completion (vs. 79% without memory and 71% with full history).
Empirical evaluation reported in the paper across three enterprise scenarios comparing three configurations: selective persistent memory, no memory, and full conversation-history persistence. Exact per-scenario sample sizes not stated in the excerpt.
Among technology firms, but not others, AI adoption is higher for firms with more employees and higher values of Tobin's q.
Reported heterogeneous associations from subgroup analyses/regressions showing positive correlation between firm size (employees), Tobin's q and adoption in technology-sector subsample but not in non-tech subsamples.
AI adoption has more than quadrupled from 5% in 2022.
Reported trend statistic comparing measured adoption rates across years (author-reported 5% baseline in 2022 and higher levels by 2025).
A further 10% were using AI in the production of goods and delivery of services.
Reported descriptive statistic from the authors' coding of 2025 SEC 10-Ks indicating AI use in production/service delivery.
In 2025, 11% of S&P 500 enterprises had AI deeply integrated into their business processes.
Reported descriptive statistic from the authors' enterprise-level adoption measure applied to S&P 500 firms in 2025 (based on SEC 10-K coding).
We develop a novel measure to assess deep AI adoption (and distinguish it from AI hype) that is based on SEC 10-K filings, where laws and regulations "prohibit companies from making materially false or misleading statements."
Methodological claim: measure constructed from corporate SEC 10-K filings; relies on legal incentives for truthful disclosure; described as novel in paper.
We study AI adoption of S&P 500 firms over the period 2016 to 2025, estimating adoption at the enterprise level.
Descriptive statement of study scope and sampling frame (S&P 500 firms, years 2016–2025); enterprise-level adoption estimation described in paper methods.
Future research priorities should include implementation science, ethical AI governance aligned with NIST AI RMF, ISO/IEC 42001, and OECD AI Principles, and SME‑specific digital resilience benchmarks to democratize data-driven decision-making in the U.S. SME sector.
Author recommendations based on the narrative review of peer‑reviewed literature (2020–2025); prescriptive statement rather than an empirical finding.
Adaptive dashboarding, cloud-based predictive models, agentic supply-chain pipelines, and machine-learning-based scenario planning are changing the operations of SMEs.
Narrative synthesis across literature (2020–2025) reported in the review; the excerpt offers no quantitative adoption rates or study counts.
There is a paradigm shift from retrospective reporting to real-time and AI‑enhanced analytics in SME business operations.
Claimed in the review based on peer‑reviewed literature (2020–2025); no aggregate metrics or counts of studies provided in the excerpt.
Small and medium-sized (SME) business organizations constitute the structural foundation of the United States economy.
Narrative statement in the review summarizing peer‑reviewed literature (2020–2025); no specific empirical sample size or citation provided in the supplied excerpt.
We release the platform, code, and dataset as a shared testbed for controlled studies of human-AI co-creation.
Statement in the paper declaring release of platform, code, and dataset; availability is presented as a factual deliverable of the project.
Prior exposure to highly creative ideas improves later performance, suggesting a 'seeding' intervention.
Experimental observation from the pilot (N = 62) that participants exposed to highly creative ideas showed improved subsequent performance; interpreted as evidence for a seeding effect.
An in-person pilot (N = 62) demonstrates the utility of the platform.
Empirical pilot study reported in the paper with sample size N = 62 (in-person).
The platform supports decomposition of performance into three typically confounded factors: participant traits, partner perceptions, and content dynamics.
Methodological claim in the paper describing platform features and intended decomposition; supported by platform design and analytic approach rather than an empirical numeric result.
We introduce a controlled, two-player extension of the Alternate Uses Test (AUT) that enables comparison of human-human and human-AI co-creation under matched interactive conditions, alongside calibrated non-interactive baselines.
Methodological description in the paper: design and implementation of a two-player AUT extension and experimental platform; no numerical sample size required for the methodological claim.
Under the Twin Transition, consumption recovers to +1.70% above baseline by 2035.
S4 scenario results from the 23-sector recursive dynamic CGE model calibrated to 2019 I-O table, simulated through 2035.
The Twin Transition (combined S4) is approximately macro-additive, producing GDP +1.06% by 2030 and +1.95% by 2035.
Simulation of combined scenario S4 in the 23-sector recursive dynamic CGE model calibrated to Vietnam 2019 I-O table, combining the S2 TFP shocks and the S3 IT investment surge.
Under Green AI, consumption rises by +1.17% by 2030 and +2.03% by 2035.
Same 23-sector recursive dynamic CGE model; scenario S2 (Green AI) as a TFP shock to heavy manufacturing and electricity.
Green AI delivers a compounding GDP dividend of +0.98% by 2030 and +1.79% by 2035.
Results from a 23-sector recursive dynamic Computable General Equilibrium (CGE) model calibrated to Vietnam's 2019 Input-Output Table and simulated through 2035; scenario S2 (Green AI) modelled as a TFP shock to heavy manufacturing and electricity.