0 cumulative citations
View corpus contextFixating on compute thresholds for LLMs misdirects AI security policy and risks turning controls into industrial policy battles; instead, weaponization should be defined by intent and observable capabilities, with live benchmarks spanning data, algorithms and compute.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
1 cumulative citations
View corpus contextU.S. AI security policy is increasingly shaped by an $\textit{LLM Mirage}$, the belief that national security risks scale in proportion to the compute used to train frontier language models. That premise fails in two ways. It miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems that use specialized data, algorithmic efficiency, and widely available hardware, while compute controls harden only a high-end perimeter. It also destabilizes regulation because, absent a settled definition of "AI weaponization," compute thresholds are easily renegotiated as domestic priorities shift, turning security policy into a proxy contest over industrial competitiveness. We analyze how the LLM Mirage took hold, propose an intent-and-capability definition of AI weaponization grounded in effects and international humanitarian law, and outline measurement infrastructure based on live benchmarks across the full AI Triad (data, algorithms, compute) for weaponization-relevant capabilities.
Summary
Main Finding
The paper diagnoses an "LLM Mirage": U.S. AI security policy has treated training compute as a legible proxy for weaponization risk. That mistake both misallocates strategic attention (hardened compute perimeters miss realistic weaponization pathways like task-specific systems using specialized data and algorithmic efficiency on commodity hardware) and politicizes regulation (compute thresholds are easy to renegotiate, concentrating agenda control with infrastructure incumbents). The authors propose replacing compute-centric governance with an effects-based, intent-and-capability definition of AI weaponization and operationalizing it via live adversarial benchmarks across the AI Triad (data, algorithms, compute).
Key Points
- LLM Mirage defined: the belief that danger scales primarily with training compute and that restricting compute suffices to prevent weaponization.
- Why compute became dominant:
- Empirical scaling narratives (power laws, Chinchilla results, emergent abilities) made compute a single-number shorthand for "frontier" capability.
- Compute is administrable: inspectable hardware, export licensing, and shipment auditing fit existing enforcement mechanisms.
- Two major failures of compute-centric governance:
- Strategic blind spots: weaponizable capabilities can be produced via specialized data, efficient algorithms, and commodity or export-compliant hardware (examples: Hunyuan-Large, DeepSeek-R1, small tactical models in Ukraine, Zhousidun dataset).
- Political instability and capture: absent a clear definition of weaponization, compute thresholds become bargaining chips in industrial/diplomatic negotiations and concentrate regulatory influence with high-capital firms.
- Measurement politics: thresholds reify categories, invite gaming (when a metric becomes a target), and privilege auditable infrastructure over demonstrated harmful effects.
- Distributional effects:
- Access and regulatory burdens are rationed by capital/endowment rather than by demonstrated risk, favoring large incumbents.
- Global South faces enforced underdevelopment and exposure to task-specific surveillance and repression tools that don’t trigger compute-based controls.
- Proposed normative shift:
- Definition: "AI weaponization is the intentional employment or objective immediate capability of an AI system to cause physical, digital, or psychological harm, to create an asymmetric advantage in conflict, or to materially facilitate such effects."
- Operationalization via live, adversarial benchmarks that evaluate capabilities across the AI Triad (data, algorithms, compute) and tie regulatory triggers to demonstrated effects/use-conditions and IHL considerations rather than raw compute.
Data & Methods
- Analytical approach: qualitative synthesis and policy analysis grounded in:
- Technical literature on scaling laws and model performance (e.g., power laws, Chinchilla, emergent abilities) to trace why compute became a salient metric.
- Political economy and measurement-literature insights on metrics-as-targets, securitization, and regulatory capture.
- Case examples and contemporary policy episodes (e.g., Biden EO framing frontier models by compute, Commerce Department export controls, the 2025 diffusion framework and its rescission, subsequent NVIDIA H200 export approvals).
- Empirical illustrations: published model deployments and datasets (Hunyuan-Large, DeepSeek-R1, Zhousidun, tactical models in Ukraine) showing weaponizable effects below frontier compute thresholds.
- Normative/legal grounding: alignment of the intent-and-capability definition with International Humanitarian Law (IHL) and an effects-first regulatory logic.
- Operational proposals: design of measurement infrastructure via continuous, adversarial benchmarks that test real-world weaponization-relevant tasks across data, algorithms, and compute dimensions (the AI Triad).
Implications for AI Economics
- Market structure & incumbency
- Compute-centric controls favor capital-intensive incumbents (chipmakers, hyperscalers), concentrating agenda power and raising barriers to entry.
- Shifting to capability/effects-based measurement would re-balance competitive advantage toward actors with specialized data and algorithmic know-how, potentially lowering the absolute capital threshold for producing operationally relevant systems.
- Industrial policy & trade
- Export controls framed around compute operate as de facto industrial policy (rationing frontier infrastructure). Replacing compute thresholds with capability-based rules would change which goods and services are trade-restricted and reduce leverage tied to semiconductor supply-chains.
- Countries that relied on hardware scarcity to maintain advantage may face decreased strategic leverage if capabilities can be produced on more broadly available hardware.
- R&D allocation & incentives
- Compute-centric governance incentivizes scale-focused investments (huge clusters, larger models). An effects-based regime and adversarial benchmarking would redirect incentives toward task-specific datasets, algorithmic efficiency, and robustness—areas with lower capital but high strategic externalities.
- Firms may invest more in hidden optimizations and specialized datasets if regulatory triggers track demonstrated capability, altering where private R&D dollars flow.
- Compliance costs and asymmetric burdens
- Current compute thresholds impose compliance costs that scale with organizational capacity, disadvantaging smaller firms, public-interest researchers, and labs in the Global South. Capability-based measurement could be more complex to audit but may reduce capital-based exclusion.
- However, designing and operating robust capability benchmarks imposes administrative and technical costs; who funds and controls these infrastructures will shape market access and governance outcomes.
- Global innovation and diffusion
- Compute-based scarcity creates enforced technological underdevelopment in many countries, encouraging reliance on Global North platforms. A move to capability-focused governance could democratize certain capabilities but also requires safeguards to prevent proliferation of harmful applications.
- Anticipate geographic shifts in comparative advantage: regions with strong data sources or domain expertise could become centers of capability even without frontier compute.
- Policy credibility and investment risk
- Durable, effects-based regulation would reduce policy volatility tied to changing political priorities over industrial competitiveness, lowering regulatory uncertainty for long-horizon investors.
- Conversely, poorly specified capability metrics could reintroduce ambiguity and gaming opportunities—making the design of transparent, adversarial benchmarks crucial for credible policy and healthy market signaling.
- Externalities & distributional risk
- The paper highlights externalities (surveillance, repression, public-safety harms) that disproportionately affect Global South populations; economics policy must account for these cross-border welfare effects when designing export and assistance programs.
- Insurance, liability, and procurement markets will need to adjust valuation of AI systems away from compute-based priors toward demonstrated behavior in context.
Overall, the paper suggests that economic policy around AI should stop treating compute as the default lever. Instead, regulators and markets should push toward capability- and effect-based measurement infrastructures that reshape incentives across R&D, trade, and competition—while explicitly accounting for distributional impacts and governance of the benchmarking infrastructure itself.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| U.S. AI security policy is increasingly shaped by an "LLM Mirage": the belief that national security risks scale in proportion to the compute used to train frontier language models. Governance And Regulation | negative | U.S. AI security policy framing (belief that risk scales with training compute) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The "LLM Mirage" miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems that use specialized data, algorithmic efficiency, and widely available hardware. Ai Safety And Ethics | negative | ability of adversaries to obtain weaponizable AI capabilities via task-specific systems |
Reading fidelity
high
Study strength
low
|
not reported
|
| Compute controls harden only a high-end perimeter (i.e., they address only the most compute-intensive frontier systems). Governance And Regulation | negative | efficacy/scope of compute-based controls |
Reading fidelity
high
Study strength
low
|
not reported
|
| Absent a settled definition of 'AI weaponization,' compute thresholds are easily renegotiated as domestic priorities shift, turning security policy into a proxy contest over industrial competitiveness. Governance And Regulation | negative | stability/character of security regulation (susceptibility to domestic political pressures) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper proposes an intent-and-capability definition of AI weaponization grounded in effects and international humanitarian law. Governance And Regulation | positive | definition/framework for identifying AI weaponization |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper outlines measurement infrastructure based on live benchmarks across the full AI Triad (data, algorithms, compute) for weaponization-relevant capabilities. Ai Safety And Ethics | positive | measurement infrastructure for weaponization-relevant AI capabilities |
Reading fidelity
high
Study strength
speculative
|
not reported
|