Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The Distributed Production System model captures how agent heterogeneity, resource constraints, communication topology, and task structure jointly determine the productivity, efficiency, and robustness of distributed systems across biology, economics, neuroscience, and computing.
Presentation of a unified theoretical model (Distributed Production System) in the paper; conceptual/mathematical development and cross-disciplinary argumentation; no empirical sample size reported.
Participants in the treatment conditions showed greater positive belief change about the AI across the session.
Pre/post measures of participant beliefs collected during the field experiment (N=388) showing larger positive shifts among those assigned to treatment conditions versus controls.
A cognitive scaffolding intervention (partnership training that reframed AI as a thought partner) was associated with higher individual document quality at the top of the distribution.
Field experiment with 388 employees comparing cognitive scaffolding to other conditions; reported improvements concentrated at the top of the individual document-quality distribution.
LLMs coordinate extremely well on similar actions.
Empirical observation from the experiment showing high coordination performance by LLMs when alignment on similar actions is the equilibrium; qualitative description in the abstract without reported quantitative metrics.
Like humans, [LLMs] regulate [action similarity] in response to coordination incentives (strategic monoculture).
Empirical claim based on experimental results comparing how humans and LLMs change similarity when incentives for coordination/divergence are manipulated. No numerical details in excerpt.
LLMs exhibit high levels of baseline similarity (primary monoculture).
Empirical observation from the experiment comparing baseline action similarity across LLM subjects (relative level described qualitatively in paper). Specific sample sizes and quantitative metrics not provided in the excerpt.
We implement a simple experimental design that cleanly separates these forces, and deploy it on human and large language model (LLM) subjects.
Methodological claim: authors report implementing an experiment that separates baseline similarity from strategic adjustments and applying it to human participants and LLM agents. No sample sizes or procedural details provided in the excerpt.
We distinguish primary algorithmic monoculture -- baseline action similarity -- from strategic algorithmic monoculture, whereby agents adjust similarity in response to incentives.
Conceptual/theoretical distinction proposed in the paper (definition and taxonomy introduced by the authors). No empirical sample size reported for this conceptual claim in the provided text.
The framework provides practical guidance for designing measurements that support identification, comparability, and efficient estimation of latent treatment effects.
Paper claims to offer practical guidance as part of the methodological framework; this is a stated contribution grounded in the theoretical framework and design recommendations (no empirical validation sample size reported).
Estimation relies on a debiasing procedure that permits valid inference even when the bridge functions are weakly identified.
Estimation approach described in the paper including a debiasing procedure and theoretical results on inference validity under weak identification of bridge functions (methodological derivations; no empirical sample size).
A design-based approach built around nonparametric bridge functions can address the noncomparability challenges; these bridge functions can be characterized and identified.
Methodological proposal with formal characterization and identification results for nonparametric bridge functions presented in the paper (theoretical proofs/derivations; no empirical sample size).
We develop a general nonparametric framework for identifying and estimating average treatment effects on latent outcomes in randomized experiments.
The paper presents a methodological contribution: a nonparametric identification and estimation framework for average treatment effects (ATE) on latent outcomes in randomized experiments. Evidence is theoretical development and formal identification arguments in the paper (no empirical sample size reported).
Overall, BDA functions as a performance amplifier, yielding higher returns for ventures well-positioned to leverage its potential.
Synthesis conclusion from empirical findings across multiple outcomes (survival, costs, sales, employee growth, financing) in the sample of German start-ups.
For high-performing BDA adopters, employee growth is even more pronounced.
Heterogeneity analysis in the paper indicating stronger employee growth among high-performing BDA adopters in the German start-up sample.
For high-performing BDA adopters, increases in sales are even more pronounced.
Paper reports heterogeneity analysis showing stronger sales effects among high-performing adopters in the German start-up sample.
Conditional on survival, BDA adopters are more likely to attract venture capital financing.
Empirical finding in the paper that surviving BDA adopters have a greater likelihood of obtaining venture capital, based on the sample of German start-ups.
Conditional on survival, BDA adopters show stronger employee growth.
Paper reports greater employee growth for surviving BDA adopters compared with non-adopters based on empirical data from German start-ups.
Conditional on survival, BDA adopters have higher sales.
Conditional (survivor) analysis reported for adopters versus non-adopters in a large sample of German start-ups.
A machine-learning research agenda is needed centered on team-level evaluation, privacy-preserving memory layers, scaffolded AI for learning, carbon-aware routing, and pro-agency workflow design.
Prescriptive recommendation in the position paper proposing specific research priorities; no empirical evaluation of these approaches is presented within the paper itself.
Rather than eliminating the office, this shift supports selective co-presence, reserving in-person time for tasks with high tacitness, high coupling, or high relational stakes (including apprenticeship, conflict repair, trust formation, and early-stage synthesis).
Theoretical/qualitative argument about task types best suited for in-person interaction; illustrated by examples (apprenticeship, conflict repair, trust formation, early-stage synthesis); no empirical task-level allocation study presented.
Capabilities that are already widely deployed—transcription, summarization, retrieval, translation, drafting, and code assistance—are the basis for this shift (with bounded agents as an amplifying but not necessary extension).
Descriptive claim citing the prevalence of specific AI capabilities in current deployments; presented as observation in the position paper rather than as a quantified adoption study.
The organizational significance of these systems is not generic automation but the accumulation of artifact capital: durable, queryable, reusable traces such as transcripts, summaries, decisions, tickets, code comments, and retrieval layers.
Argumentative claim in the paper describing a conceptual mechanism ('artifact capital') by which foundation-model features create reusable organizational artifacts; no empirical measurement of artifact capital provided.
The foundation-model stack (NL interaction, multimodal capture, long context, retrieval, transcription, translation, bounded tool use) changes the coordination economics that previously favored daily in-person co-presence.
Conceptual claim supported by descriptions of foundation-model capabilities and their potential to create durable, queryable artifacts; no empirical test or measured coordination-costs reported.
Remote-capable knowledge work should default to AI-enabled flexibility because the workflow-integrated foundation-model stack changes the coordination economics that once favored daily co-presence.
Normative argument in the position paper based on conceptual analysis of coordination economics and the claimed effects of foundation-model features; no empirical sample or quantitative study reported.
Preliminary corroboration is provided by a companion production automation system with eleven operating lanes and 2,132 classified tickets.
Reported companion system operational statistics in the paper (11 lanes, 2,132 tickets).
When iteration was permitted, the final success rate for the structured interactions reached 91.5% (183 of 200).
Reported final success counts/rate in the paper for structured interactions (183 of 200).
Among structured interactions, 110 of 200 were accepted on first pass.
Reported counts in the paper for the structured-interaction group (110 accepted of 200 structured interactions).
Structured context assembly was associated with an improvement in first-pass acceptance from 32% to 55%.
Observational comparison reported in the paper (baseline vs. structured first-pass acceptance rates are given as 32% and 55%).
Structured context assembly was associated with a reduction from 3.8 to 2.0 average iteration cycles per task.
Observational comparison reported in the paper (structured vs. baseline interactions); the paper states the 3.8 to 2.0 cycle figures.
The paper applies formal models from reliability engineering and information theory as post hoc interpretive lenses on context quality.
Paper text claiming the application of these formal models for interpretation.
Context Engineering applies a staged four-phase pipeline (Reviewer to Design to Builder to Auditor).
Methodological description in the paper listing the four pipeline phases.
Context Engineering defines a five-role context package structure (Authority, Exemplar, Constraint, Rubric, Metadata).
Explicit specification in the paper of the five-role package components.
This paper introduces Context Engineering, a structured methodology for assembling, declaring, and sequencing the complete informational payload that accompanies a prompt to an AI tool.
Methodological description in the paper (definition and presentation of the Context Engineering approach).
The review integrates fragmented literature into a cohesive framework and offers implications for managers and policymakers to pursue more balanced, inclusive, and context-sensitive AI adoption strategies.
Author-stated contribution of the review based on synthesis of the 40 included studies; normative recommendations derived from the review.
Generative AI adoption is associated with mixed employee perceptions: some studies report increased efficiency and higher job satisfaction.
Aggregate finding from included studies in the review that report positive employee-reported outcomes (efficiency, satisfaction).
There is consistent evidence of productivity improvements from generative AI in workplace settings, driven by task automation, decision support, and knowledge augmentation.
Synthesis of findings across the 40 included empirical and conceptual studies (review-level conclusion summarising multiple studies reporting productivity effects).
The realization of the positive effect of big data applications on markups depends on the synergistic support of various complementary resources.
Authors conclude—based on model analysis and empirical heterogeneity tests—that complementary resources (organizational, technological, environmental) are necessary for big data applications to translate into higher markups; details and sample sizes not provided in the summary.
Improving production efficiency is a key channel through which big data applications contribute to higher price markups.
Mechanism analysis in the paper identifies improved production efficiency as the second key channel linking big data applications to increased markups; supported by the heterogeneous firm model and empirical tests on firm-level data (sample size not reported).
Promotion of product innovation is a key channel through which big data applications contribute to higher price markups.
Mechanism analysis in the paper identifies product innovation as one of two key channels linking big data applications to increased markups; supported by the model and empirical mechanism tests using micro-level firm data (sample size not reported).
Big data applications significantly enhance firms' price markups.
Paper constructs a heterogeneous firm model with variable markups and conducts empirical tests using micro-level firm data; reported result states a significant positive effect of big data applications on firms' price markups. (Sample size not reported in the provided summary.)
Ireland’s high levels of educational attainment offer a strong foundation for benefiting from AI adoption, but targeted educational support (especially for older workers or those with lower formal qualifications) and investment in lifelong learning and retraining will be essential.
Policy assessment based on Ireland's workforce characteristics and the report's scenario findings about which groups face disruption; presented as a recommendation/interpretation.
Increases in returns to capital as a result of AI adoption, while modest in percentage terms, benefit households at the very top of the income distribution, where the vast majority of Ireland’s capital income is concentrated.
Simulated changes in returns to capital combined with income distribution data showing concentration of capital income among top households; reported in the report.
For those who remain in work, AI is expected to increase productivity. We estimate that workers who are not displaced may see modest but broadly shared wage gains.
Scenario assumptions and international evidence on productivity effects of AI, incorporated into the report's simulations of wages for non-displaced workers.
We present a gaze-grounded multimodal LLM assistant that uses egocentric video with gaze overlays to identify likely points of difficulty and target follow-up retrospective assistance.
System description and implementation presented in the paper: an assistant combining egocentric video and gaze overlays to detect potential user difficulties and provide retrospective help.
Gaze-aware LLM assistants can reason about cognitive needs to improve cognitive outcomes of users.
Authors' synthesis and interpretation of controlled-study results (n=36) showing improved recall, perceived accuracy/personalization, and more efficient interactions under the gaze-aware condition.
Users spoke significantly fewer words with the gaze-aware assistant, indicating more efficient interactions.
Behavioral measure recorded during the controlled study (n=36): word count of user speech in gaze-aware vs text-only conditions; authors report a statistically significant reduction in words spoken in the gaze-aware condition.
The gaze-aware assistant significantly improved people's ability to recall information.
Controlled study (n=36) comparing recall performance between gaze-aware and text-only assistant conditions; authors report a statistically significant improvement in recall for the gaze-aware condition.
Compared to a conventional LLM assistant, the gaze-aware assistant was rated as significantly more personalized in its assessments of users' reading behavior.
Between-subjects controlled study (n=36) using user ratings of personalization for the gaze-aware vs text-only assistant; authors report a statistically significant increase in perceived personalization for the gaze-aware condition.
Compared to a conventional LLM assistant, the gaze-aware assistant was rated as significantly more accurate in its assessments of users' reading behavior.
Between-subjects controlled study (n=36) comparing user ratings of the gaze-aware assistant vs a text-only LLM; authors report a statistically significant difference in perceived accuracy of assessments.
Exploitative innovation is directly associated with long-term competitive performance.
PLS-SEM analysis of survey data from 104 Portuguese B2B managers showing a significant direct path from exploitative innovation to performance.