The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

To secure both national interests and AI development, China should adopt a tiered regulatory regime for cross‑border generative AI training data; drawing on EU rights‑focused and U.S. market/security approaches, the paper recommends differentiated data preconditions, an independent regulator with multilevel coordination, and international articulation of China's data governance.

Risk Drivers and Regulatory Pathways for Cross-Border Flows of Generative AI Training Data
NianZhu Qiu · August 01, 2026 · Law and Humanities
openalex descriptive n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. NianZhu Qiu provider ID

Semantic Scholar

Latest observation:

  1. NianZhu Qiu provider ID
The paper analyzes legal risks from cross-border flows of generative AI training data, contrasts EU and U.S. regulatory models, and proposes a tiered Chinese governance framework that balances security, development, and international coordination.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As a representative form of new quality productive forces, generative artificial intelligence (AI) depends fundamentally on the global circulation of data, creating a deeply coupled and mutually reinforcing relationship between model development and cross-border data flows. The scale of training data, the transnational spillover of associated risks, and the difficulty of effective oversight have nevertheless generated multidimensional security concerns involving national security, corporate interests, and personal privacy. Using a comparative analysis of the regulatory approaches adopted by the European Union and the United States, this article examines the structural sources of these risks and evaluates the practical limitations of existing governance arrangements. Grounded in China's domestic conditions and the holistic approach to national security, it proposes a tiered regulatory architecture based on three elements: differentiated preconditions for data protection and circulation, an independent regulator supported by multilevel coordination, and the international articulation of a Chinese approach to data governance. This framework seeks to balance security and development while contributing to a more coherent global regime for cross-border flows of generative AI training data.

Summary

Main Finding

Cross-border flows of generative AI training data create large-scale, transboundary security, economic, and privacy risks driven by data scale, spillovers, and fragmented governance; a workable response for China is a tiered, classified regulatory architecture (differentiated data preconditions, an independent regulator with multilevel coordination, and proactive international engagement) that balances national security and AI development while reducing regulatory uncertainty and contributing to global rule-making.

Key Points

  • Risk drivers
    • Scale: training modern generative models requires massive, heterogeneous datasets from multiple jurisdictions, increasing volume and velocity of cross-border transfers.
    • Spillover effects: noncompliance or embedded biases can propagate internationally via models, affecting economic, political, and military domains.
    • Regulatory difficulty: jurisdictional fragmentation, divergent norms (sovereignty vs. free data flow), and the dynamic data needs of ML complicate oversight.
  • Structural legal sources of risk
    • Global inequalities in access to training data create competitive asymmetries and constrain policy choices for many states.
    • Divergent national rules (localization, approval procedures, liability allocation) produce overlapping obligations, regulatory-selection paradoxes, and compliance uncertainty.
    • Data mobility undermines effective sovereign control; localization can produce silos and unintended vulnerabilities.
  • Practical governance challenges
    • Equity: unregulated flows risk “data colonialism,” concentrating value in a few multinationals and reducing developing countries to raw-data suppliers.
    • National security: multimodal datasets may contain strategic information (infrastructure, geolocation, health/genetics) that, if transferred, increase exposure.
    • Privacy and re-identification: generative models intensify re-identification risks and complicate rights such as the right to be forgotten.
  • Comparative models
    • European Union: rights-oriented, stringent risk control (risk classification in AI Act, GDPR adequacy, SCCs/BCRs), high compliance costs but strong privacy and strategic-data protections.
    • United States: market-oriented pragmatism combining decentralized law with executive-national-security instruments (CLOUD Act, Executive Orders) that both facilitate allied transfers and block access for “countries of concern.”
  • Proposed Chinese approach
    • Leverage existing legal matrix (National Security Law; Cybersecurity, Data Security, Personal Information Protection Laws; Measures and Interim AI rules).
    • Adopt classified/graded protection of training data and proportionate transfer conditions (free flow for low-risk data; restricted for sensitive categories).
    • Create an independent regulator supported by multilevel coordination and full-process auditability.
    • Pursue international interoperability while articulating a distinct Chinese governance position.

Data & Methods

  • Methodology: qualitative comparative legal and policy analysis; normative institutional design.
  • Sources analyzed: primary legal instruments and policies (EU AI Act and GDPR mechanisms; U.S. legislative/ executive measures including CLOUD Act and relevant Executive Orders; Chinese Cybersecurity Law, Data Security Law, Personal Information Protection Law, Measures for Cross-Border Data Transfers, Interim Measures for Generative AI), plus secondary literature on AI governance, security, and economic distributional effects.
  • No empirical datasets or econometric analysis—argumentation builds from legal texts, regulatory practice, and techno-legal characteristics of generative AI.

Implications for AI Economics

  • Market structure and competition
    • Access to cross-border training data is a competitive input; restrictive regimes raise incumbents’ advantage where they control large datasets and infrastructure, potentially entrenching digital monopolies.
    • A tiered Chinese system that enables low-risk flows while restricting sensitive data could help domestic firms access diverse training inputs and capture more AI value, but will raise compliance costs.
  • Innovation and productivity
    • Broad, frictionless data access lowers marginal costs of model improvement and accelerates innovation; conversely, fragmentation/localization raises costs, slows iteration, and may reduce global diffusion of productivity gains.
    • Targeted regimes (classification/grading) can balance innovation incentives with protection of strategic assets if implemented with clarity and predictability.
  • Trade, investment, and global value chains
    • Differing regimes (EU strictness, U.S. strategic filtering, China’s tiered control) increase transaction costs for cross-border data-dependent services and may fragment markets, altering FDI patterns in cloud and AI services.
    • Adequacy frameworks, SCCs/BCRs, and interoperability arrangements are economically important mechanisms that reduce frictions and enable trade in data-intensive services.
  • Distributional effects and digital development
    • Unregulated flows risk value capture by advanced economies/tech firms; regulated, reciprocity-aware policies can protect domestic data value and foster local AI ecosystems.
    • However, tighter controls may impede smaller firms’ access to global datasets, raising entry costs and necessitating supporting public infrastructure (data commons, compute).
  • Policy uncertainty and risk premia
    • Regulatory fragmentation and rapid rule changes raise compliance risk and regulatory uncertainty, increasing cost of capital for AI ventures and potentially slowing investment.
    • Transparent, rule-based tiering with clear criteria reduces informational and compliance frictions, lowering risk premia and supporting longer-term investment in domestic AI capabilities.

Overall, regulatory design around cross-border training-data flows will materially shape comparative advantages in AI, distribution of rents from data, innovation trajectories, and the international structure of AI markets. The paper’s proposed tiered, interoperable Chinese framework aims to preserve development potential while managing those economic risks, but will require careful calibration to avoid excessive compliance costs or unintended protectionism.

Assessment

Paper Typedescriptive Evidence Strengthn/a — The paper is a legal and policy analysis with no empirical identification strategy or causal inference; it does not present empirical tests or quantitative evidence of effects. Methods Rigorlow — The article uses comparative legal analysis and normative argumentation rather than systematic empirical methods; while the legal reasoning is coherent, there is no formal methodology, robustness checks, or empirical validation of the proposed regulatory framework. SampleQualitative analysis of statutory and regulatory texts, official documents, and policy instruments (e.g., EU AI Act/GDPR adequacy framework, U.S. CLOUD Act and executive orders, Chinese Cybersecurity Law, Data Security Law, Personal Information Protection Law, and recent Chinese AI interim measures); comparative review of regulatory approaches and normative proposal for a tiered Chinese governance architecture. Themesgovernance innovation inequality GeneralizabilityRecommendations are specific to China's legal and institutional context and may not transfer directly to other jurisdictions., Normative proposals are untested; practical effectiveness and economic impacts are not empirically evaluated., Focuses on generative AI training-data flows; may not generalize to other AI models or data types without modification., Assumes certain levels of state capacity and political will that vary across countries.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Cross-border flows of generative-AI training data are characterized by immense scale, pronounced transnational spillover effects, and exceptional regulatory difficulty. Governance And Regulation negative Regulatory manageability and security risks associated with cross-border training-data flows
Reading fidelity high
Study strength low
not reported
0.09
No unified international legal framework currently governs cross-border flows of generative-AI training data; countries primarily rely on domestic legislation. Governance And Regulation negative International regulatory coordination
Reading fidelity high
Study strength low
not reported
0.09
Stronger data-localization policies may disperse computing resources, create data silos, and enlarge security vulnerabilities, causing legislative objectives and practical outcomes to diverge. Organizational Efficiency negative Security and operational consequences of data-localization policies
Reading fidelity high
Study strength low
not reported
0.09
Unregulated cross-border training-data flows may concentrate data resources in a small number of multinational technology corporations and leave developing countries as suppliers of training corpora without sharing proportionately in AI-generated value. Inequality negative Distribution of digital resources and value across countries and firms
Reading fidelity high
Study strength speculative
not reported
0.03
Unrestricted cross-border flows of AI training data can expose domestic infrastructure and security safeguards to attack, creating political, economic, and military risks. Ai Safety And Ethics negative National-security and critical-infrastructure vulnerability
Reading fidelity high
Study strength speculative
not reported
0.03
The European Union uses a rights-oriented, risk-based model for cross-border AI training-data governance, including risk classification, adequacy requirements, and contractual mechanisms such as Standard Contractual Clauses and Binding Corporate Rules. Governance And Regulation positive Regulatory control and privacy protection for cross-border training-data transfers
Reading fidelity high
Study strength medium
not reported
0.18
The EU's stringent regulatory model increases compliance costs while establishing strong market-access barriers and extending EU regulatory influence over cross-border strategic-data flows. Market Structure mixed Compliance costs and regulatory influence over cross-border data markets
Reading fidelity high
Study strength medium
not reported
0.18
The United States combines reduced data-transfer barriers with allied countries and interventionist restrictions on designated geopolitical competitors, treating cross-border AI training-data security as an issue of national security and data sovereignty. Governance And Regulation mixed Strategic regulation and geopolitical allocation of cross-border training-data access
Reading fidelity high
Study strength medium
not reported
0.18
China's existing governance framework provides differentiated compliance pathways for cross-border data transfers, including security assessments, regulatory filing, and risk reporting. Regulatory Compliance positive Regulatory compliance pathways for cross-border training-data transfers
Reading fidelity high
Study strength medium
not reported
0.18
China should adopt a classified and graded regulatory strategy under which cross-border risk assessments and management measures vary according to the type, importance, and security level of generative-AI training data. Governance And Regulation positive Risk-based regulatory effectiveness and proportionality of data controls
Reading fidelity high
Study strength speculative
not reported
0.03

Notes