0 cumulative citations
View corpus contextCombining AI with supercomputing yields disproportionately novel and highly cited science, but the computational advantages are consolidating in a handful of regions—chiefly the United States and China—raising the prospect of widening global disparities in discovery.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
1 cumulative citations
View corpus contextAbstract Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000–2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output). The future of discovery will depend not only on advances in algorithms and computing power, but also on enacting policies that democratise these capabilities across the global scientific ecosystem.
Summary
Main Finding
Research that explicitly combines AI and high-performance computing (AI+HPC) is strongly associated with frontier scientific discovery: AI+HPC papers are more likely to be both novel (introducing new terminology later reused) and highly cited (top 1%). This synergy yields substantially larger gains than AI-only or HPC-only work. At the same time, access to the necessary compute and expertise is highly concentrated geographically and institutionally (notably the US and China, with the EU27 competitive in aggregate), raising risks of a more stratified global science system.
Key Points
- Scope and scale
- Dataset: >5.3 million Scopus-indexed publications across 27 fields, 2000–2024.
- AI publications grew from ~40k (2000) to >600k (2024).
- HPC acknowledgements appear in ~1% of papers in 2024 (≈6× growth since 2015).
- By 2024, ~16% of HPC-tagged papers incorporated AI; ~1% of AI papers explicitly acknowledged HPC.
- Frontier outcomes
- Frontier discovery defined as both highly novel (introducing a term reused later) and highly impactful (top 1% citations).
- Baseline probability of introducing a novel term ≈ 5% (papers using neither AI nor HPC).
- Predicted novelty: AI-only or HPC-only papers ≈ 7–8%; AI+HPC papers ≈ 16% (roughly three times baseline and greater than additive).
- Example: In Biochemistry, Genetics & Molecular Biology, 5% of AI+HPC papers are in the top 1% most-cited (≈5× baseline).
- Convergence type
- AI+HPC work splits into “AI-development” (building/improving models) and “AI-use” (applying AI to domain problems). Since ~2017, growth is concentrated in applied AI-use leveraging HPC.
- Concentration of compute and publications
- Supercomputing capacity and AI-specific compute are highly concentrated: estimates cited US ≈75% of global AI supercomputing power (April 2025), China ≈15%; private sector capacity growing relative to academia/government.
- Country- and institution-level inequality in AI/HPC publications has increased (Mean Log Deviation measures reported).
- Positive association between cumulative computational capacity (FLOPs) and number of breakthrough AI+HPC publications.
- Robustness and limits
- Results robust across sensitivity analyses (alternative novelty measures, keyword definitions, exclusion of CS, etc.).
- Limitations: observational design (no causal identification of compute → discovery), conservative HPC detection via acknowledgements (recall ~62–70% against full-text LLM validation), relatively small sample of AI+HPC papers (wider CIs).
Data & Methods
- Bibliometric corpus
- Source: Scopus metadata for 2000–2024; 5.3M+ records; 27 ASJC fields.
- Extracted: authors, affiliations, abstracts, citations, acknowledgements, (where available) full text.
- Identification of AI and HPC
- AI: curated list of 203 keywords (machine learning, deep learning, GenAI, etc.) applied to titles/abstracts/keywords; sensitivity checks vs narrower keyword sets.
- HPC: mined acknowledgements for explicit mentions (e.g., “HPC”, “supercomputer”); validated via case studies and LLM full-text checks.
- AI+HPC: intersection of the two sets.
- Characterisation of AI+HPC work
- LLM-based classifier (applied to full texts where available) to separate AI-development vs AI-use contributions.
- Outcome measures
- Impact: top 1% most-cited papers (field- and year-normalized).
- Novelty: papers introducing a new word/phrase (term) that is subsequently reused in later literature.
- Statistical analysis
- Descriptive statistics and field-level comparisons.
- Regression models predicting novelty (and other robustness checks), controlling for affiliation, prior citations, year, and other confounders.
- Inequality measures: Mean Log Deviation on Top500 FLOPs and publication counts over time.
- Validation and robustness
- LLM validation of HPC detection; alternative novelty definitions; exclusion tests; multiple sensitivity analyses reported in Supplementary Materials.
Implications for AI Economics
- Compute as a scarce, high-return input
- Evidence shows strong complementarities between AI and HPC that amplify both novelty and impact. In economic terms, compute behaves like a scarce, high-marginal-product input with non-linear returns when combined with AI expertise.
- Implication: production functions for scientific output should include compute (FLOPs, access to HPC) and AI-human capital as distinct inputs, with interaction terms to capture super-additivity.
- Concentration, rents, and market power
- Large shares of AI/HPC capacity concentrated in a few countries and private firms imply potential for concentrated economic rents and market power in both science and downstream innovation (commercialization).
- Policy and antitrust economists should examine how compute concentration affects pricing of compute services, access for public research, and knowledge diffusion.
- Inequality in scientific capacity and growth externalities
- Geographic concentration may generate persistent comparative advantages and path-dependence (agglomeration economies), increasing returns to regions with early compute investment and depriving others of spillovers.
- International and regional policy interventions (compute subsidies, shared HPC centers, data/infrastructure sharing) can be evaluated for efficiency and equity trade-offs.
- R&D allocation and distortion risks
- If AI+HPC-equipped labs disproportionately attract funding and attention, research portfolios may skew toward problems amenable to compute-intensive methods (and commercializable outputs), potentially undervaluing other socially important but less compute-intensive research.
- Funders should consider counterfactuals: how much frontier science is contingent on compute access vs other inputs?
- Labor and skill complementarities
- Demand for AI/HPC skills across scientific labor markets will rise, altering returns to human capital (higher wages for compute-savvy scientists) and career incentives; this can deepen intra- and inter-country skill gaps.
- Measurement and policy evaluation opportunities
- The paper demonstrates feasible metrics for measuring compute-access effects (acknowledgements, FLOPs, Top500) and frontier output (novelty via new-term reuse). These can be used to evaluate policies (e.g., EuroHPC, national supercomputer investments) via quasi-experimental designs.
- Suggested research agendas for AI economics
- Causal identification: exploit exogenous shocks (new supercomputer installations, funding changes, scheduled decommissioning, or cross-border hardware export controls) to estimate causal effects of compute on novelty, citations, patents, and commercialization.
- Compute in R&D production functions: estimate elasticities of output (publications, high-impact discoveries, patents) with respect to compute and its interaction with AI skill capital.
- Market structure and pricing: study pricing dynamics of cloud/HPC providers, barriers to entry, and effects on non-profit and public research institutions.
- Welfare and distributional analysis: quantify social returns vs private capture of AI+HPC-enabled discoveries; optimal subsidy or sharing mechanisms.
- Spillovers and diffusion: measure how breakthroughs enabled by concentrated compute diffuse to other firms/countries, and what institutional arrangements accelerate diffusion.
- Labor markets: quantify returns to AI/HPC skills and the effect on academic career trajectories and inequality.
- Policy experiments: evaluate shared infrastructure models (national HPC centers, time-allocation policies, open-access compute credits) for efficiency and equity outcomes.
Limitations to bear in mind when using the paper as evidence - Correlational results: the paper shows strong associations but cannot fully rule out selection (elite labs adopting compute earlier). - Measurement gaps: HPC detection via acknowledgements is conservative (recall ~62–70%) and may undercount compute use; AI keyword approaches may misclassify some interdisciplinary work. - Evolving technology: compute costs, architectures, and software ecosystems evolve rapidly; findings (e.g., US/China compute shares) are snapshots that can shift with policy or private investment.
Concise takeaway AI and supercomputing together markedly increase the likelihood of frontier discoveries, but compute and expertise are increasingly concentrated. For AI economists this implies large returns to compute, potential for persistent geographic and institutional advantages, and urgent needs for policy design and empirical causal work to understand welfare and distributional consequences.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Research combining AI and HPC is more likely to introduce novel ideas than either conventional work or research using AI or HPC in isolation. Research Productivity | positive | introduction of novel ideas (idea novelty measure) |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| Research combining AI and HPC is more likely to achieve top-cited status than either conventional work or research using AI or HPC in isolation. Research Productivity | positive | top-cited status (citation impact / top percentile citations) |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| The computational synergy between AI and HPC is most pronounced at the scientific frontier. Research Productivity | positive | magnitude of AI+HPC advantage in novelty and high citation outcomes at the scientific frontier |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| There are growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions. Inequality | negative | geographic concentration of access to supercomputing resources and AI expertise (access disparities) |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| These concentrated regions are dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output. Adoption Rate | mixed | regional share and competitiveness of AI+HPC output (high-impact publications) |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| The combined impact of AI and HPC correlates with scientific breakthroughs. Research Productivity | positive | scientific breakthroughs proxied by novelty and top-cited metrics |
Reading fidelity
high
Study strength
medium
|
n=5000000
|
| Policy action to democratise AI and HPC capabilities across the global scientific ecosystem is necessary for the future of discovery. Governance And Regulation | positive | policy effectiveness in democratizing capabilities (recommendation, not empirically tested in paper) |
Reading fidelity
high
Study strength
speculative
|
n=5000000
|