Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We introduce AgroVG, a multi-source benchmark that formulates agricultural grounding as generalized set prediction: given an image and a referring expression, a model must return all matching target instances or abstain when no target is present.
Paper contribution: description of a new benchmark and its task formulation (benchmark construction and formalization).
Evaluating agricultural visual grounding therefore requires jointly testing localization accuracy, target-set completeness, and existence-aware abstention.
Methodological assertion in the paper motivating the benchmark design (conceptual requirement for evaluation metrics and protocols).
Visual grounding is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, and targeted harvesting.
Framing / motivation statement in the paper abstract/introduction (conceptual argument linking visual grounding capability to downstream agri-applications).
Enterprise capability adaptation serves as the key support for implementing intelligent international marketing models.
Conclusion from the paper's review and content analysis of literature (2010–2025); presented as a synthesized enabling factor rather than empirically quantified effect.
Mainstream innovation models include data-driven precision marketing, AI-powered cross-border CRM, intelligent omnichannel integration, and cross-cultural intelligent localization marketing.
Summary from the paper's systematic review and content analysis of core literature (2010–2025); descriptive synthesis, no primary experimental sample size reported.
New theoretical frameworks have emerged: data-driven precision marketing theory, nonlinear customer journey reconstruction theory, cross-border intelligent value co-creation theory, and global intelligent marketing ecosystem theory.
Identified via the paper's systematic review and content analysis of literature from 2010–2025; presented as conceptual/theoretical developments rather than quantified empirical effects.
Intelligent technologies have increased international marketing ROI by 12%–25%.
Mixed-method systematic review and content analysis of core literature sources from 2010 to 2025 (as reported in the paper). No primary dataset or sample size reported for this quantified range.
Structured production-process management and size are significant predictors of AI adoption.
Regression/associational analysis from the Census Bureau survey showing that measures of structured production-process management and establishment size predict reported AI use; sample ~28,500 establishments.
Deployed FLUID increases Active Hours by +0.05%.
Reported online metric improvement from production experiments/deployment as stated in the paper. No statistical significance, confidence intervals, or sample sizes provided in the excerpt.
Deployed FLUID increases Cold-Start Room Views by +2.05%.
Reported online metric improvement from production experiments/deployment as stated in the paper. No statistical significance, confidence intervals, or sample sizes provided in the excerpt.
Deployed FLUID delivers an online gain of +0.55% Quality Watch Duration.
Reported online metric improvement from production experiments/deployment as stated in the paper. No statistical significance, confidence intervals, or sample sizes provided in the excerpt.
FLUID was deployed on industrial livestreaming recommenders with a cross-platform combined user base of over one billion globally.
Authors' deployment statement in the paper indicating production rollout across industrial recommenders and noting a combined user base (statement of scope/scale). No A/B sample sizes reported in the excerpt.
FLUID uses a late-fusion, ID-free design that injects slice-level and room-level LUCID as independent tokens, stabilized by a staged warmup under online incremental training.
Methodological/system design description in the paper specifying late-fusion ID-free architecture, token injection strategy, and staged warmup for online incremental training.
FLUID couples a cross-domain multimodal encoder, jointly trained on short videos and livestreams, to produce discrete hierarchical codes (LUCID).
Methodological description in paper: joint training on short videos and livestreams to generate discrete hierarchical codes named LUCID. This is a system/design claim from the methods section.
FLUID is the first framework to fully retire the candidate-side item ID from a production-scale livestreaming ranker.
Authors' claim of novelty and system capability; supported in-document by description of FLUID's ID-free design and production deployment note. No independent verification provided in the excerpt.
Live-agent performance depends on objective tracking, execution conversion, cost, and runtime reliability, supporting evaluation of LLMs as components in bounded workflows rather than as isolated benchmark respondents.
Synthesis of experimental results (cross-provider differences in end-to-end play, planner bakeoff, and trace analyses) that link specific mechanisms (objective tracking, execution conversion, cost, runtime reliability) to performance.
In a replicated 32-game cross-provider championship under frozen rules, gemini-3.1-pro-preview won 20 of 32 games against gpt-5.1, claude-opus-4-7, and kimi-k2.6, and the pooled winner distribution differs strongly from an equal-strength null (p approx 1.5 x 10^-5).
Empirical tournament experiment: 32 games played under frozen rules across four provider models; reported win counts and a statistical test vs an equal-strength null yielding p ≈ 1.5×10^-5.
The deep integration of the digital and real economies and the accumulation of human capital are fundamental drivers of sound and rapid development of the overall economy.
Theoretical framing and empirical emphasis in the paper asserting the importance of human capital (digital talent) and digital-real economy integration for economic growth; supported by the paper’s cross-regional empirical analysis linking digitalization and talent to growth outcomes.
Digital talent agglomeration and industrial digitalization are important drivers of regional economic growth.
Overall empirical results from cross-provincial/regional analysis in China reported in the paper, which link measures of digital talent concentration and industrial digitalization to regional economic growth outcomes.
In the Yangtze River Delta region, digital talent agglomeration and industrial digitalization have achieved a positive and interactive relation that promotes regional economic growth.
Regional-case empirical analysis focused on the Yangtze River Delta showing a positive interaction between talent agglomeration and industrial digitalization associated with higher regional economic growth (reported in the paper). Specific sample size for the region is not stated in the excerpt.
Focusing on observation instead of prediction, and governance rather than control, complements existing alignment and safety practices while preserving human judgment, institutional choice, and long-term wellbeing.
Normative argument presented in the paper linking observational monitoring to governance objectives; no empirical evaluation provided.
Interpretable, aggregate behavioral signals (as described) support human-in-the-loop interpretation and enable earlier awareness of when AI use patterns may be drifting from creative augmentation toward automation pressure, authority substitution, or unintended displacement of human agency.
Conceptual claim about intended use of monitoring signals; no empirical test or sample presented.
A system-level framework for externalized behavioral monitoring should treat generative AI systems as participants in socio-technical ecosystems rather than static tools, emphasizing interpretable, aggregate behavioral signals such as shifts in output velocity, semantic and structural reuse, persistence of synthetic roles, and cross-context propagation.
Proposed conceptual framework and list of candidate behavioral signals in the paper (design/specification, no empirical validation).
Post-deployment observability is a foundation for well-being-aligned human–AI co-evolution.
Conceptual argument and system-level framework presented in the paper (no empirical study or sample reported).
The findings carry significant implications for entrepreneurs, policymakers, and educators seeking to leverage AI as a driver of inclusive and sustainable entrepreneurial success in urban India.
Authors' stated implications in the discussion and conclusion sections, derived from thematic findings across the 16 interviews.
An entrepreneur's mindset—specifically cognitive openness, risk tolerance, and iterative experimentation—is the strongest predictor of successful AI adoption outcomes, superseding firm size, sector, and financial capacity.
Cross-cutting finding from thematic analysis of the 16 interview transcripts indicating recurring emphasis on mindset attributes as drivers of successful adoption; comparative qualitative assessment across interviewees suggested these factors mattered more than firm size, sector, or finances.
Overall, AI adoption produces measurable benefits in operational efficiency, strategic decision-making, and customer personalisation among the entrepreneurs studied.
Synthesis of interview findings/themes from the 16-case qualitative study; authors state AI adoption 'produces measurable benefits' across these domains based on participant reports.
AI acts as a competitive equaliser among entrepreneurs in Delhi/NCR.
Theme 'AI as a Competitive Equaliser' produced by thematic analysis of the 16 interviews; participants reported that AI lowered barriers and allowed smaller firms to compete more effectively.
AI adoption transforms customer experience by enabling greater personalisation.
Theme 'Customer Experience Transformation' from thematic analysis of interviews (n=16); entrepreneurs described AI-driven personalisation and improved customer interactions.
AI adoption improves strategic decision-making and market intelligence among entrepreneurs.
One of five thematic findings ('AI-Enabled Decision Making and Market Intelligence') derived from thematic analysis of 16 interviews; participants reported using AI for market insights and better decisions.
AI functions as an operational accelerator for entrepreneurs, producing benefits in operational efficiency.
Thematic analysis of interview data (n=16) generated a theme labelled 'AI as an Operational Accelerator' reporting interviewee accounts of operational efficiency gains.
This study integrates observed GenAI uses into a coherent, processual view of growth hacking by developing first-order concepts, second-order themes and three aggregate dimensions mapped onto a seven-stage growth pipeline.
Methodological claim supported by the study's adopted approach: Gioia methodology applied to 17 semi-structured interviews with founders/growth leaders (nine startups), plus secondary sources.
Generative AI reallocates human attention from asset production to problem framing, inference quality and organizational learning across the seven stages of the growth pipeline.
Interview-derived themes (17 interviews across nine startups) and process mapping of GenAI uses onto the seven-stage growth pipeline.
Generative AI acts as a data orchestrator that automates cleaning, cohorting, variance checks and knowledge capture, tightening feedback loops and institutionalizing learning.
Findings derived from 17 semi-structured interviews with founders and growth leaders across nine startups, supported by secondary sources and Gioia-style thematic analysis.
Generative AI serves as a cognitive sparring partner that reduces bounded rationality and groupthink via premortems, counter-arguments and stakeholder role-plays while preserving human judgment.
Same qualitative data set of 17 interviews across nine startups, with Gioia-method coding producing first-order concepts and themes describing AI-mediated decision practices.
Generative AI functions as an experimentation accelerator, lowering the marginal cost of variation and compressing the idea-to-test cycle, enabling parallel selections of controlled tests.
Exploratory multiple-case qualitative study using 17 semi-structured interviews with founders and growth leaders across nine startups, plus secondary sources; analysis via the Gioia methodology to derive themes mapped onto a seven-stage growth pipeline.
Large-scale validation in a production code completion environment shows Echo increased the acceptance rate from 25.7% to 35.7%.
Reported result from 'large-scale validation' in a production code completion environment as stated in the paper's abstract; no sample size, statistical tests, or additional experimental details provided in the excerpt.
User-driven refinement sequences distill agents' flawed proposals into high-quality training signals.
Conceptual/empirical claim in the paper that user refinements produce verified solutions which serve as high-quality signals; supported by the paper's later validation claim but no separate sample size or statistical detail provided in the excerpt.
Echo is a generalized framework that operationalizes the transition from raw experience to learnable knowledge by echoing environmental feedback into the training loop for model optimization.
Methodological contribution described by the authors (framework description); no implementation details or quantitative validation given in the excerpt besides later mention of validation.
Widespread deployment of AI agents provides low-cost access to massive streams of real-world experience data.
Stated observation in the paper; no quantitative deployment statistics or sample sizes provided in the excerpt.
Continuous learning from 'experience data' (interactions between agents and their environments) promises to transcend the scalability and knowledge limitations of static human data.
Conceptual claim in the paper proposing continuous learning from experience data as a solution; no empirical details provided in the excerpt.
There is a session-level carryover effect: a participant's prior AI use leads to further AI adoption and entrenches their miscalibration about time savings.
Observed analyses across sessions in the three pre-registered user studies (combined N = 2691) showing that prior within-session AI use predicts subsequent AI adoption and stronger miscalibration.
People display 'efficiency-gain illusions': they overestimate how much time and effort savings AI use provides.
Same three pre-registered user studies (combined N = 2691) that measured participants' perceived time/effort savings from AI versus actual measured time/effort.
People frequently choose to use AI even when doing so is inefficient (i.e., provides no meaningful time or effort savings).
Three pre-registered user studies reported in the paper (combined N = 2691) measuring participants' choices to use AI on cognitively simple tasks and comparing those choices to measured time/effort savings.
Future research should focus on empirical assessments of the economic ramifications of artificial intelligence, particularly regarding productivity enhancement, labour market restructuring, and equitable income distribution.
Recommendation in the paper's discussion/conclusion based on identified gaps and the theoretical model (no empirical study presented to support specific magnitudes).
Regulatory bodies should ensure access to data, support platform markets, and promote that artificial intelligence redistributes wealth among the owners of capital, data and labour.
Normative recommendation grounded in the paper's theoretical-legal model and comparative policy discussion (method: deductive/inductive reasoning; no empirical intervention or evaluation).
The European Union has established a comprehensive legal and regulatory framework for the digital economy and artificial intelligence, including rules on platform usage, digital goods liability, data protection (GDPR), and AI.
Comparative legal review of EU regulations and statutes described in the paper (method: comparative approach).
The rise of digital technologies and artificial intelligence will dramatically improve the way existing economic systems function.
Theoretical synthesis and comparative legal analysis presented in the paper; no empirical data or sample reported (methodology: inductive and deductive reasoning, comparative approach).
The result is a shared vocabulary for practitioners building hybrid systems, an analytical lens for researchers studying combination patterns, and a starting point for evaluators interested in the full quality of human-AI decision-making rather than accuracy alone.
Authors' stated contributions/anticipated utility of their framework (conceptual claim about the expected usefulness of their mapping).
Closing the synergy gap requires explicit engagement with a wider design space.
Prescriptive conclusion from the authors advocating broader design engagement (conceptual recommendation based on their framework).