The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Combining AI with supercomputing yields disproportionately novel and highly cited science, but the computational advantages are consolidating in a handful of regions—chiefly the United States and China—raising the prospect of widening global disparities in discovery.

Scientific discovery in the age of AI and supercomputing
Stefano Bianchini, Aldo Geuna, Fazliddin Shermatov · July 25, 2026 · Scientific Reports
openalex correlational medium evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Stefano Bianchini provider ID
  2. Aldo Geuna provider ID
  3. Fazliddin Shermatov provider ID

Semantic Scholar

Latest observation:

  1. Stefano Bianchini provider ID
  2. A. Geuna provider ID
  3. Fazliddin Shermatov provider ID
Research that combines artificial intelligence and high-performance computing is more likely to introduce novel ideas and become highly cited, while access to these computational capabilities is increasingly concentrated in a few regions (notably the US and China).

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000–2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output). The future of discovery will depend not only on advances in algorithms and computing power, but also on enacting policies that democratise these capabilities across the global scientific ecosystem.

Summary

Main Finding

Research that explicitly combines AI and high-performance computing (AI+HPC) is strongly associated with frontier scientific discovery: AI+HPC papers are more likely to be both novel (introducing new terminology later reused) and highly cited (top 1%). This synergy yields substantially larger gains than AI-only or HPC-only work. At the same time, access to the necessary compute and expertise is highly concentrated geographically and institutionally (notably the US and China, with the EU27 competitive in aggregate), raising risks of a more stratified global science system.

Key Points

  • Scope and scale
    • Dataset: >5.3 million Scopus-indexed publications across 27 fields, 2000–2024.
    • AI publications grew from ~40k (2000) to >600k (2024).
    • HPC acknowledgements appear in ~1% of papers in 2024 (≈6× growth since 2015).
    • By 2024, ~16% of HPC-tagged papers incorporated AI; ~1% of AI papers explicitly acknowledged HPC.
  • Frontier outcomes
    • Frontier discovery defined as both highly novel (introducing a term reused later) and highly impactful (top 1% citations).
    • Baseline probability of introducing a novel term ≈ 5% (papers using neither AI nor HPC).
    • Predicted novelty: AI-only or HPC-only papers ≈ 7–8%; AI+HPC papers ≈ 16% (roughly three times baseline and greater than additive).
    • Example: In Biochemistry, Genetics & Molecular Biology, 5% of AI+HPC papers are in the top 1% most-cited (≈5× baseline).
  • Convergence type
    • AI+HPC work splits into “AI-development” (building/improving models) and “AI-use” (applying AI to domain problems). Since ~2017, growth is concentrated in applied AI-use leveraging HPC.
  • Concentration of compute and publications
    • Supercomputing capacity and AI-specific compute are highly concentrated: estimates cited US ≈75% of global AI supercomputing power (April 2025), China ≈15%; private sector capacity growing relative to academia/government.
    • Country- and institution-level inequality in AI/HPC publications has increased (Mean Log Deviation measures reported).
    • Positive association between cumulative computational capacity (FLOPs) and number of breakthrough AI+HPC publications.
  • Robustness and limits
    • Results robust across sensitivity analyses (alternative novelty measures, keyword definitions, exclusion of CS, etc.).
    • Limitations: observational design (no causal identification of compute → discovery), conservative HPC detection via acknowledgements (recall ~62–70% against full-text LLM validation), relatively small sample of AI+HPC papers (wider CIs).

Data & Methods

  • Bibliometric corpus
    • Source: Scopus metadata for 2000–2024; 5.3M+ records; 27 ASJC fields.
    • Extracted: authors, affiliations, abstracts, citations, acknowledgements, (where available) full text.
  • Identification of AI and HPC
    • AI: curated list of 203 keywords (machine learning, deep learning, GenAI, etc.) applied to titles/abstracts/keywords; sensitivity checks vs narrower keyword sets.
    • HPC: mined acknowledgements for explicit mentions (e.g., “HPC”, “supercomputer”); validated via case studies and LLM full-text checks.
    • AI+HPC: intersection of the two sets.
  • Characterisation of AI+HPC work
    • LLM-based classifier (applied to full texts where available) to separate AI-development vs AI-use contributions.
  • Outcome measures
    • Impact: top 1% most-cited papers (field- and year-normalized).
    • Novelty: papers introducing a new word/phrase (term) that is subsequently reused in later literature.
  • Statistical analysis
    • Descriptive statistics and field-level comparisons.
    • Regression models predicting novelty (and other robustness checks), controlling for affiliation, prior citations, year, and other confounders.
    • Inequality measures: Mean Log Deviation on Top500 FLOPs and publication counts over time.
  • Validation and robustness
    • LLM validation of HPC detection; alternative novelty definitions; exclusion tests; multiple sensitivity analyses reported in Supplementary Materials.

Implications for AI Economics

  • Compute as a scarce, high-return input
    • Evidence shows strong complementarities between AI and HPC that amplify both novelty and impact. In economic terms, compute behaves like a scarce, high-marginal-product input with non-linear returns when combined with AI expertise.
    • Implication: production functions for scientific output should include compute (FLOPs, access to HPC) and AI-human capital as distinct inputs, with interaction terms to capture super-additivity.
  • Concentration, rents, and market power
    • Large shares of AI/HPC capacity concentrated in a few countries and private firms imply potential for concentrated economic rents and market power in both science and downstream innovation (commercialization).
    • Policy and antitrust economists should examine how compute concentration affects pricing of compute services, access for public research, and knowledge diffusion.
  • Inequality in scientific capacity and growth externalities
    • Geographic concentration may generate persistent comparative advantages and path-dependence (agglomeration economies), increasing returns to regions with early compute investment and depriving others of spillovers.
    • International and regional policy interventions (compute subsidies, shared HPC centers, data/infrastructure sharing) can be evaluated for efficiency and equity trade-offs.
  • R&D allocation and distortion risks
    • If AI+HPC-equipped labs disproportionately attract funding and attention, research portfolios may skew toward problems amenable to compute-intensive methods (and commercializable outputs), potentially undervaluing other socially important but less compute-intensive research.
    • Funders should consider counterfactuals: how much frontier science is contingent on compute access vs other inputs?
  • Labor and skill complementarities
    • Demand for AI/HPC skills across scientific labor markets will rise, altering returns to human capital (higher wages for compute-savvy scientists) and career incentives; this can deepen intra- and inter-country skill gaps.
  • Measurement and policy evaluation opportunities
    • The paper demonstrates feasible metrics for measuring compute-access effects (acknowledgements, FLOPs, Top500) and frontier output (novelty via new-term reuse). These can be used to evaluate policies (e.g., EuroHPC, national supercomputer investments) via quasi-experimental designs.
  • Suggested research agendas for AI economics
    • Causal identification: exploit exogenous shocks (new supercomputer installations, funding changes, scheduled decommissioning, or cross-border hardware export controls) to estimate causal effects of compute on novelty, citations, patents, and commercialization.
    • Compute in R&D production functions: estimate elasticities of output (publications, high-impact discoveries, patents) with respect to compute and its interaction with AI skill capital.
    • Market structure and pricing: study pricing dynamics of cloud/HPC providers, barriers to entry, and effects on non-profit and public research institutions.
    • Welfare and distributional analysis: quantify social returns vs private capture of AI+HPC-enabled discoveries; optimal subsidy or sharing mechanisms.
    • Spillovers and diffusion: measure how breakthroughs enabled by concentrated compute diffuse to other firms/countries, and what institutional arrangements accelerate diffusion.
    • Labor markets: quantify returns to AI/HPC skills and the effect on academic career trajectories and inequality.
    • Policy experiments: evaluate shared infrastructure models (national HPC centers, time-allocation policies, open-access compute credits) for efficiency and equity outcomes.

Limitations to bear in mind when using the paper as evidence - Correlational results: the paper shows strong associations but cannot fully rule out selection (elite labs adopting compute earlier). - Measurement gaps: HPC detection via acknowledgements is conservative (recall ~62–70%) and may undercount compute use; AI keyword approaches may misclassify some interdisciplinary work. - Evolving technology: compute costs, architectures, and software ecosystems evolve rapidly; findings (e.g., US/China compute shares) are snapshots that can shift with policy or private investment.

Concise takeaway AI and supercomputing together markedly increase the likelihood of frontier discoveries, but compute and expertise are increasingly concentrated. For AI economists this implies large returns to compute, potential for persistent geographic and institutional advantages, and urgent needs for policy design and empirical causal work to understand welfare and distributional consequences.

Assessment

Paper Typecorrelational Evidence Strengthmedium — Very large, multi-field dataset (5+ million papers) and consistent correlational patterns give credible descriptive evidence that AI+HPC work is associated with greater novelty and citation impact, but the analysis is observational, subject to measurement error in identifying AI/HPC use, field- and time-varying citation dynamics, and potential confounders and reverse causality that prevent strong causal claims. Methods Rigormedium — Scale, multi-field coverage and use of metadata are methodological strengths that support robustness checks and heterogeneity analysis, but the abstract does not describe identification strategies that address endogeneity (e.g., instrumenting, natural experiments, or causal discontinuities), nor does it detail how novelty and resource use are measured or how confounding is handled, leaving substantive risk of bias. SampleMetadata from more than five million scientific publications spanning 2000–2024 across 27 scientific fields; includes bibliometric indicators (novelty metrics and citation counts) and institution/country attribution for mapping AI and supercomputing usage and output. Themesinnovation productivity inequality adoption human_ai_collab IdentificationAssociational analysis of publication metadata: papers are classified as AI-use, HPC-use, AI+HPC, or neither (likely via keywords, methods, acknowledgements or resource mentions) and outcomes (novelty metrics, top-cited status) are compared across these groups, probably using regressions with controls for field, year and observable covariates and descriptive country/institution counts; no experimental or quasi-experimental causal source is reported in the abstract. GeneralizabilityObservational publication data may not capture all uses of AI or HPC (underreporting or inconsistent acknowledgment across fields)., Citation-based outcomes vary widely by field and time and may reflect network and visibility effects rather than scientific quality alone., Findings reflect published research only and omit unpublished work, preprints not indexed similarly, or industrial R&D., Country/institution comparisons may be biased by differences in publication norms, language, and database coverage., Associational results do not identify causal mechanisms, limiting inference to correlational statements about impact on discovery.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Research combining AI and HPC is more likely to introduce novel ideas than either conventional work or research using AI or HPC in isolation. Research Productivity positive introduction of novel ideas (idea novelty measure)
Reading fidelity high
Study strength medium
n=5000000
0.3
Research combining AI and HPC is more likely to achieve top-cited status than either conventional work or research using AI or HPC in isolation. Research Productivity positive top-cited status (citation impact / top percentile citations)
Reading fidelity high
Study strength medium
n=5000000
0.3
The computational synergy between AI and HPC is most pronounced at the scientific frontier. Research Productivity positive magnitude of AI+HPC advantage in novelty and high citation outcomes at the scientific frontier
Reading fidelity high
Study strength medium
n=5000000
0.3
There are growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions. Inequality negative geographic concentration of access to supercomputing resources and AI expertise (access disparities)
Reading fidelity high
Study strength medium
n=5000000
0.3
These concentrated regions are dominated by the United States and China, though the EU27 aggregate maintains high competitiveness in combined AI+HPC output. Adoption Rate mixed regional share and competitiveness of AI+HPC output (high-impact publications)
Reading fidelity high
Study strength medium
n=5000000
0.3
The combined impact of AI and HPC correlates with scientific breakthroughs. Research Productivity positive scientific breakthroughs proxied by novelty and top-cited metrics
Reading fidelity high
Study strength medium
n=5000000
0.3
Policy action to democratise AI and HPC capabilities across the global scientific ecosystem is necessary for the future of discovery. Governance And Regulation positive policy effectiveness in democratizing capabilities (recommendation, not empirically tested in paper)
Reading fidelity high
Study strength speculative
n=5000000
0.05

Notes