Evidence (10501 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filtered →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Innovation
Remove filter
Science has a positive local effect on co-located digital and artistic activity while exerting a negative regional backwash effect that draws creative capacity away from neighbouring districts.
Spatial Durbin model estimating within-district and between-district effects for digital technology, science, and arts segments.
Creative agglomeration in Slovakia is conditional on reaching segment-specific density thresholds and is shaped by asymmetric cross-district spillovers and industrial legacy.
District-level analysis using segment-specific location quotients, population density, manufacturing specialization, a spatial Durbin model, random-forest simulations, and breakpoint tests.
No single capability-sourcing route dominates under all conditions; the headline results depend on the corresponding model mechanisms.
Mechanism-knockout experiments that switch acquisition friction, absorption, open-weight cost reductions, and substitution on or off.
Waiting is relatively safe for a firm whose existing product cannot be replaced by the emerging technology, but it causes substantial value loss for a firm whose product is directly substitutable by that technology.
Agent-based model with demand-side substitution eroding the legacy revenue of non-adapters according to product exposure; stated as Proposition P4.
A one-time reduction in the cost of building generative-AI capability revives the build route primarily before the leading design has stabilized; after the dominant design is established, the effect is substantially weaker.
Agent-based model experiment varying the timing of an open-weight-style reduction in build costs relative to dominant-design formation; stated as Proposition P3.
The route by which incumbents enter generative AI depends on contractibility: firms tend to partner when the capability can be rented through an API and tend to absorb when the capability is too tacit to rent.
History-friendly agent-based simulation of incumbent sourcing choices across build, partner, acquire, absorb, and wait routes; stated as Proposition P1 and reported as a headline finding.
Platform regulations such as transparency requirements, auditability, and algorithmic impact assessments can change the incentives and equilibrium outcomes associated with algorithmic design in news markets.
Theoretical policy implication; the article proposes evaluating such effects but does not estimate them.
High-quality journalism has public-good characteristics, while algorithmic amplification of low-quality content can generate negative externalities.
Normative and economic framing in the discussion of externalities and public goods; no welfare estimate or empirical comparison is provided.
Algorithmic ranking and recommendation can allocate attention in ways that create distributional externalities for news producers and affect demand for different types of content.
Theoretical implication for platform economics and attention markets; the article does not quantify the externalities or estimate effects.
Accountability for public knowledge is distributed across journalists, platforms, and regulators rather than being assigned to a single actor.
Theoretical account of accountability within the contested epistemic space; no observed cases or measured accountability outcomes are reported.
Editorial choices and algorithmic amplification jointly shape what audiences see, while editorial standards, algorithmic signals, and regulatory legitimacy jointly shape what audiences trust.
Conceptual decomposition of the framework into visibility and credibility mechanisms; not supported by an empirical sample.
The intersection of editorial, algorithmic, and regulatory logics forms a contested epistemic space in which visibility, credibility, and accountability are negotiated and may be stabilized or destabilized.
Theoretical framework identifying the intersection and its effects on public knowledge; no empirical validation is presented.
Journalism's epistemic authority emerges from the interaction of editorial, algorithmic, and regulatory logics rather than being produced solely within newsrooms.
Conceptual synthesis in the article's main finding; no empirical sample or causal test is reported.
The interaction of the AI Act and GDPR creates synergies in governance and accountability but can also compound compliance burdens, producing trade-offs between reduced harms and trust on one hand and efficiency and innovation on the other.
Synthesis of the doctrinal analysis and the paper's discussion of compliance costs, trust, innovation, and regulatory interaction; no quantitative trade-off estimate is provided.
The combined regulatory regimes may induce firms to relocate activities, partition product lines, or maintain dual compliance tracks, thereby affecting the location of data processing and AI development.
Analytical inference from the interaction of extraterritorial obligations, compliance costs, and regulatory differences; no firm-level relocation data are reported.
EU rules may diffuse globally as de facto standards because firms adopt uniform EU-aligned practices to preserve market access, potentially reducing regulatory fragmentation while exporting EU norms.
Comparative legal analysis of extraterritoriality and market-access incentives, interpreted through the regulatory-governance and regulatory-capitalism literature.
Governance-by-design may shift innovation toward safer and more auditable systems, while potentially slowing time-to-market.
Theoretical and interpretive analysis of how embedded compliance requirements may affect product design and innovation incentives; no quantitative innovation or time-to-market estimate is provided.
The relationships between AI adoption orientation, entrepreneurial intention, entrepreneurial behaviour, and business performance vary by venture stage, supporting the characterization of AI adoption as a stage-contingent capability.
The study used measurement-invariance testing and PLS-SEM multi-group analysis to compare 108 new and 86 established women entrepreneurs.
AI perspective is primarily associated with entrepreneurial outcomes among established women entrepreneurs rather than equally across both venture stages.
Multi-group PLS-SEM analysis of new and established women entrepreneurs; the abstract reports that AI perspective becomes significant primarily in established ventures.
Rapid technological change in diagnostics, pharmaceuticals, and digital health is shifting which countries produce key health inputs.
Qualitative policy analysis based on observed technological diffusion and changes in production capacity.
Digital transformation has stronger positive associations with substantive carbon disclosure components—carbon-related business practices, carbon governance, and carbon performance—than with the disclosure carrier or reporting format/channel.
Component-level regressions separating substantive content elements from the disclosure-carrier component of the multidimensional CIDQ index.
Following the COVID-19 shock, European airlines converged in some network dimensions while remaining persistently differentiated in others.
The post-pandemic longitudinal comparison evaluates OT-based network dimensions across the 32-airline sample.
Relative to a 2016 baseline, the longitudinal analysis identifies diverse adjustment trajectories across airlines.
Airline-year distributions are compared with a 2016 baseline over the 2016–2025 observation window.
In an application to 32 European airlines observed from 2016 through 2025, airline similarities were multidimensional and varied by carrier type and network feature.
The empirical analysis uses scheduled passenger-flight data for 32 European airlines over the 2016–2025 period and compares carriers across multiple OT dimensions.
The AI Publication Footprint is an absolute cumulative Scopus publication count and therefore serves as a proxy for knowledge-production capacity rather than a normalized measure of research intensity or efficiency.
Measurement definition and methodological caveat; publication counts were not normalized per capita or per researcher.
The cluster analysis identifies four country types arranged along a gradient from high digital maturity and high national readiness to low readiness and a low AI publication footprint.
Cluster analysis of country profiles combining AI publication footprint and national AI readiness.
Coding was the only reported domain showing mild forgetting for Thomson-1.0-Large relative to its base model, while general mathematical and abstract reasoning remained within the Qwen performance level.
Authors' interpretation of cross-domain results, including coding scores of 39.9% for Thomson-1.0-Large versus 40.9% for Qwen3.5-397B and reasoning scores of 68.4% versus 66.8%.
In the 150-request Claude-MCP command benchmark, 73.3% of requests were executed directly, 22.7% required clarification, and 4.0% requested unavailable operations.
Command-profile analysis of the deployed interactive Claude-MCP agent.
Four of the five primitives were built and running in private pilots, while supply-chain functionality was built as separate tooling but had not yet been integrated into the request path.
Authors' implementation-status report; no pilot sample size or outcome measurements are reported.
The spatial spillover of AI changes over time from an initial negative inter-city factor-siphoning effect to a later positive green-technology radiation effect.
A time-varying spatial Durbin model is applied to the 282-city panel to estimate dynamic cross-city spillovers.
Digital-green synergy is more strongly associated with strategic, breakthrough, and exploitative innovation than with substantive, incremental, and exploratory innovation.
The study compares associations across different innovation types using panel-level empirical analysis.
For the U.S. FEMA benchmark, PPE outperformed expert baselines on the Socioeconomic and Composite risk indicators, achieving R² of 66.9% versus 61.1%, while performing on par with expert benchmarks across the broader environmental-target suite.
Spatial regression evaluation of 21 FEMA environmental risk scores using approximately 84,000 census-tract observations.
Regulatory and institutional responses, including intellectual-property law, data-protection regimes, and antitrust enforcement, interact with entrepreneurial actions to shape whether AI property regimes are private, public, or hybrid.
Conceptual institutional analysis and proposed AI-governance research agenda.
In AI markets, entrepreneurial choices such as open-source versus proprietary development, licensing, and platform access controls will shape the excludability of AI outputs and the market structure of AI industries.
Application of the paper's conceptual framework to AI commercialization; presented as an implication or prediction rather than an empirical finding.
Property-rights entrepreneurship can produce new markets, shifts in firm boundaries, reallocation of rents, and institutional change.
Conceptual analysis of the potential outcomes associated with product, asset, and resource rights.
The Station outperformed AlphaEvolve on three of the seven AlphaEvolve problems that did not produce results novel relative to the prior literature, matched it on two, and underperformed it on two.
Portfolio-level comparison across all 12 evaluated problems, categorized by relative performance.
The value difference between productive evaluation of an active policy and restart evaluation of a frontier policy can be exactly decomposed into a restart-policy-quality term plus a productive-path-reuse term.
Equation 4.6 adds and subtracts restart evaluation of the active policy to produce an exact algebraic decomposition.
Organizational selection can shift toward platform-mediated commercialization before that route generates greater total surplus than incumbent-mediated collaboration.
Model comparison between evolutionary viability, determined by organizational-selection payoffs, and productive efficiency, determined by total surplus; incumbent outside-option requirements create a wedge between the two thresholds.
Supply chain digitalization functions as a contingent dynamic capability rather than a universally accessible technological fix.
This conclusion is based on the reported asymmetric subgroup effects and the paper's interpretation that resource constraints limit firms' ability to benefit from digitalization.
The proposed architecture is intended to balance blockchain trust and verifiability with the flexibility, scalability, and expressiveness of centralized metadata services.
Architectural rationale for separating immutable blockchain records from dynamic centralized services; no measured latency, cost, scalability, or security results are reported.
The study introduces reputational adjacency, describing reputational pressure on a firm caused by a competitor’s technological success rather than by failure of the focal firm.
Conceptual development based on the DeepSeek R1 case and observed stakeholder judgment shifts.
CEOs most often used an acknowledge–reframe response pattern when they responded to the disruption.
Directed content analysis of coded CEO communicative moves across the sampled communications.
Research attention is highly uneven across simulator capabilities: controllability, interaction, and stability receive substantially more attention than asset construction, the physics engine, and state feedback.
Cumulative paper counts for each capability in the 200-paper corpus, categorized by principal contribution.
World models have achieved functional substitution for interaction and controllability in specific scenarios, but they remain short of traditional simulators in formal physical-law guarantees, structured state feedback, and reproducibility of long-horizon evolution.
Comparative capability analysis of 200 papers spanning latent-dynamics, autoregressive, diffusion, JEPA, explicit 3D/4D, and occupancy-centric methods.
The study concludes that geospatial transferability depends not only on model design but also on intrinsic geographic differences between source and target regions.
Interpretation of the transfer experiments and regression analysis relating geographic domain shift to transfer performance.
Cross-region transfer performance exhibits substantial spatial heterogeneity and asymmetry across source-target geographic pairs.
Results summarized in the abstract and contribution statement, based on cross-region transfer evaluation and linear mixed-effects regression.
Geospatial transferability varies substantially across target geographic areas.
Cross-region transfer experiments using human mobility generation models trained in one geographic region and evaluated in unseen regions.
A positive Moran spatial shift indicates that geographic covariates are more spatially clustered in the target domain than in the source domain, whereas a negative shift indicates that they are more spatially dispersed.
Interpretation of the defined Moran spatial shift, calculated as the target-domain Moran's I minus the source-domain Moran's I.
The proposed geographic domain-shift framework captures two complementary forms of difference between source and target regions: feature-distribution differences through mutual-information shift and spatial-structure differences through Moran spatial shift.
Definitions and methodological description of the two proposed domain-shift metrics.
In one ablation study, removing the language-model component from several LLM-based forecasters left accuracy unchanged or improved it.
A component ablation study cited by the review, comparing complete LLM-based forecasters with versions lacking the language-model component.