The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Fixating on compute thresholds for LLMs misdirects AI security policy and risks turning controls into industrial policy battles; instead, weaponization should be defined by intent and observable capabilities, with live benchmarks spanning data, algorithms and compute.

The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls
Gupta, Ritwik, Reddie, Andrew W. · January 08, 2026 · ArXiv.org
openalex commentary n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Gupta, Ritwik provider ID
  2. Reddie, Andrew W. provider ID

Semantic Scholar

Latest observation:

  1. Ritwik Gupta provider ID
  2. A. Reddie provider ID
Focusing AI security policy on LLM training compute creates a misleading 'LLM Mirage'—security risks arise from task-specific models, data, and algorithms too—so regulation should use an intent-and-capability definition of weaponization and live benchmarks across data, algorithms, and compute.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

U.S. AI security policy is increasingly shaped by an $\textit{LLM Mirage}$, the belief that national security risks scale in proportion to the compute used to train frontier language models. That premise fails in two ways. It miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems that use specialized data, algorithmic efficiency, and widely available hardware, while compute controls harden only a high-end perimeter. It also destabilizes regulation because, absent a settled definition of "AI weaponization," compute thresholds are easily renegotiated as domestic priorities shift, turning security policy into a proxy contest over industrial competitiveness. We analyze how the LLM Mirage took hold, propose an intent-and-capability definition of AI weaponization grounded in effects and international humanitarian law, and outline measurement infrastructure based on live benchmarks across the full AI Triad (data, algorithms, compute) for weaponization-relevant capabilities.

Summary

Main Finding

The paper diagnoses an "LLM Mirage": U.S. AI security policy has treated training compute as a legible proxy for weaponization risk. That mistake both misallocates strategic attention (hardened compute perimeters miss realistic weaponization pathways like task-specific systems using specialized data and algorithmic efficiency on commodity hardware) and politicizes regulation (compute thresholds are easy to renegotiate, concentrating agenda control with infrastructure incumbents). The authors propose replacing compute-centric governance with an effects-based, intent-and-capability definition of AI weaponization and operationalizing it via live adversarial benchmarks across the AI Triad (data, algorithms, compute).

Key Points

  • LLM Mirage defined: the belief that danger scales primarily with training compute and that restricting compute suffices to prevent weaponization.
  • Why compute became dominant:
    • Empirical scaling narratives (power laws, Chinchilla results, emergent abilities) made compute a single-number shorthand for "frontier" capability.
    • Compute is administrable: inspectable hardware, export licensing, and shipment auditing fit existing enforcement mechanisms.
  • Two major failures of compute-centric governance:
  • Strategic blind spots: weaponizable capabilities can be produced via specialized data, efficient algorithms, and commodity or export-compliant hardware (examples: Hunyuan-Large, DeepSeek-R1, small tactical models in Ukraine, Zhousidun dataset).
  • Political instability and capture: absent a clear definition of weaponization, compute thresholds become bargaining chips in industrial/diplomatic negotiations and concentrate regulatory influence with high-capital firms.
  • Measurement politics: thresholds reify categories, invite gaming (when a metric becomes a target), and privilege auditable infrastructure over demonstrated harmful effects.
  • Distributional effects:
    • Access and regulatory burdens are rationed by capital/endowment rather than by demonstrated risk, favoring large incumbents.
    • Global South faces enforced underdevelopment and exposure to task-specific surveillance and repression tools that don’t trigger compute-based controls.
  • Proposed normative shift:
    • Definition: "AI weaponization is the intentional employment or objective immediate capability of an AI system to cause physical, digital, or psychological harm, to create an asymmetric advantage in conflict, or to materially facilitate such effects."
    • Operationalization via live, adversarial benchmarks that evaluate capabilities across the AI Triad (data, algorithms, compute) and tie regulatory triggers to demonstrated effects/use-conditions and IHL considerations rather than raw compute.

Data & Methods

  • Analytical approach: qualitative synthesis and policy analysis grounded in:
    • Technical literature on scaling laws and model performance (e.g., power laws, Chinchilla, emergent abilities) to trace why compute became a salient metric.
    • Political economy and measurement-literature insights on metrics-as-targets, securitization, and regulatory capture.
    • Case examples and contemporary policy episodes (e.g., Biden EO framing frontier models by compute, Commerce Department export controls, the 2025 diffusion framework and its rescission, subsequent NVIDIA H200 export approvals).
    • Empirical illustrations: published model deployments and datasets (Hunyuan-Large, DeepSeek-R1, Zhousidun, tactical models in Ukraine) showing weaponizable effects below frontier compute thresholds.
  • Normative/legal grounding: alignment of the intent-and-capability definition with International Humanitarian Law (IHL) and an effects-first regulatory logic.
  • Operational proposals: design of measurement infrastructure via continuous, adversarial benchmarks that test real-world weaponization-relevant tasks across data, algorithms, and compute dimensions (the AI Triad).

Implications for AI Economics

  • Market structure & incumbency
    • Compute-centric controls favor capital-intensive incumbents (chipmakers, hyperscalers), concentrating agenda power and raising barriers to entry.
    • Shifting to capability/effects-based measurement would re-balance competitive advantage toward actors with specialized data and algorithmic know-how, potentially lowering the absolute capital threshold for producing operationally relevant systems.
  • Industrial policy & trade
    • Export controls framed around compute operate as de facto industrial policy (rationing frontier infrastructure). Replacing compute thresholds with capability-based rules would change which goods and services are trade-restricted and reduce leverage tied to semiconductor supply-chains.
    • Countries that relied on hardware scarcity to maintain advantage may face decreased strategic leverage if capabilities can be produced on more broadly available hardware.
  • R&D allocation & incentives
    • Compute-centric governance incentivizes scale-focused investments (huge clusters, larger models). An effects-based regime and adversarial benchmarking would redirect incentives toward task-specific datasets, algorithmic efficiency, and robustness—areas with lower capital but high strategic externalities.
    • Firms may invest more in hidden optimizations and specialized datasets if regulatory triggers track demonstrated capability, altering where private R&D dollars flow.
  • Compliance costs and asymmetric burdens
    • Current compute thresholds impose compliance costs that scale with organizational capacity, disadvantaging smaller firms, public-interest researchers, and labs in the Global South. Capability-based measurement could be more complex to audit but may reduce capital-based exclusion.
    • However, designing and operating robust capability benchmarks imposes administrative and technical costs; who funds and controls these infrastructures will shape market access and governance outcomes.
  • Global innovation and diffusion
    • Compute-based scarcity creates enforced technological underdevelopment in many countries, encouraging reliance on Global North platforms. A move to capability-focused governance could democratize certain capabilities but also requires safeguards to prevent proliferation of harmful applications.
    • Anticipate geographic shifts in comparative advantage: regions with strong data sources or domain expertise could become centers of capability even without frontier compute.
  • Policy credibility and investment risk
    • Durable, effects-based regulation would reduce policy volatility tied to changing political priorities over industrial competitiveness, lowering regulatory uncertainty for long-horizon investors.
    • Conversely, poorly specified capability metrics could reintroduce ambiguity and gaming opportunities—making the design of transparent, adversarial benchmarks crucial for credible policy and healthy market signaling.
  • Externalities & distributional risk
    • The paper highlights externalities (surveillance, repression, public-safety harms) that disproportionately affect Global South populations; economics policy must account for these cross-border welfare effects when designing export and assistance programs.
    • Insurance, liability, and procurement markets will need to adjust valuation of AI systems away from compute-based priors toward demonstrated behavior in context.

Overall, the paper suggests that economic policy around AI should stop treating compute as the default lever. Instead, regulators and markets should push toward capability- and effect-based measurement infrastructures that reshape incentives across R&D, trade, and competition—while explicitly accounting for distributional impacts and governance of the benchmarking infrastructure itself.

Assessment

Paper Typecommentary Evidence Strengthn/a — This is a conceptual/policy analysis rather than an empirical study; it advances arguments and a proposed framework but does not present causal identification or new quantitative evidence. Methods Rigormedium — The paper offers a structured critique, grounds its proposed intent-and-capability definition in international humanitarian law, and outlines concrete measurement infrastructure; however, it lacks empirical validation, formal modeling, or implementation results that would elevate rigor to high. SampleNo empirical sample; the paper synthesizes policy documents, historical precedents, technical literature on large language models and compute, and legal frameworks (international humanitarian law) to motivate its argument and measurement proposals. Themesgovernance innovation adoption GeneralizabilityUS-centric policy framing may not map to other jurisdictions' regulatory or legal contexts, Conceptual proposals require operationalization and empirical testing before general application, Rapidly evolving ML techniques could change the relevance of compute-centric vs. task-specific pathways, Adversary capabilities, incentives, and access to data/hardware vary across countries and non-state actors, Feasibility of the proposed benchmarks and measurement infrastructure depends on industry cooperation and international coordination

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
U.S. AI security policy is increasingly shaped by an "LLM Mirage": the belief that national security risks scale in proportion to the compute used to train frontier language models. Governance And Regulation negative U.S. AI security policy framing (belief that risk scales with training compute)
Reading fidelity high
Study strength medium
not reported
0.06
The "LLM Mirage" miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems that use specialized data, algorithmic efficiency, and widely available hardware. Ai Safety And Ethics negative ability of adversaries to obtain weaponizable AI capabilities via task-specific systems
Reading fidelity high
Study strength low
not reported
0.03
Compute controls harden only a high-end perimeter (i.e., they address only the most compute-intensive frontier systems). Governance And Regulation negative efficacy/scope of compute-based controls
Reading fidelity high
Study strength low
not reported
0.03
Absent a settled definition of 'AI weaponization,' compute thresholds are easily renegotiated as domestic priorities shift, turning security policy into a proxy contest over industrial competitiveness. Governance And Regulation negative stability/character of security regulation (susceptibility to domestic political pressures)
Reading fidelity high
Study strength medium
not reported
0.06
The paper proposes an intent-and-capability definition of AI weaponization grounded in effects and international humanitarian law. Governance And Regulation positive definition/framework for identifying AI weaponization
Reading fidelity high
Study strength speculative
not reported
0.01
The paper outlines measurement infrastructure based on live benchmarks across the full AI Triad (data, algorithms, compute) for weaponization-relevant capabilities. Ai Safety And Ethics positive measurement infrastructure for weaponization-relevant AI capabilities
Reading fidelity high
Study strength speculative
not reported
0.01

Notes