The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

U.S. law tends to tolerate web scraping for AI training if data are publicly accessible, while the EU's GDPR makes scraping personal data legally risky unless a clear lawful basis exists, creating regulatory friction for AI developers and prompting calls for reform.

Data scraping for AI model training and data privacy: a comparative analysis of United States and European Union Data Privacy Laws
Olumide Timothy Ajayi, Chukwuemezie Charles Emejuo, Francis Aondongu Wayo, Samuel Esezoobo · September 09, 2026 · Humanities and Social Sciences Communications
openalex descriptive n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Olumide Timothy Ajayi provider ID
  2. Chukwuemezie Charles Emejuo provider ID
  3. Francis Aondongu Wayo provider ID
  4. Samuel Esezoobo provider ID
Comparative doctrinal analysis finds U.S. law generally permits scraping of publicly accessible web data for AI training, while EU GDPR places strict lawful-basis requirements on processing personal data—making large-scale scraping of personal data legally fraught absent specific legal grounds.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract Artificial Intelligence (AI) has made data not only the choicest but also a very susceptible commodity. Even before the recent boom in AI, data was regarded as the new oil of the digital era. However, AI has amplified this due to its dependence on data for training AI models. In sourcing data for AI training, the internet is a rich resource with abundant data. Consequently, companies engage in massive data scraping to provide sufficient data for effective AI model training. Meanwhile, scraping often involves unauthorized access to information on websites or internet platforms, thereby raising legal concerns, including privacy-related issues. In this note, the authors explore the intersection between AI model training and data scraping using a doctrinal research method, further interrogating the lawfulness of scraping under extant data privacy laws in the United States (U.S.) and the European Union (EU). Findings show that in the U.S., scraping data for AI may be permissible as long as the scraped data is publicly accessible. On the other hand, data scraping of publicly available personal data violates the EU data protection laws, unless it satisfies the lawful basis requirements of Article 6 of the General Data Protection Regulation (GDPR). Whilst data scraping is not per se illegal under EU law, it is almost unjustifiable in light of the stringent requirements of Article 6 of the GDPR. Despite these stringent requirements, it is worth noting that in practice, the enforcement of the provisions of GDPR by the Data Protection Authority (DPA) has not been uniform across member states, considering the lenient approach in some jurisdictions. Regardless of the differences in the U.S. and EU jurisdictions, further analysis shows that data scraping generally undermines data privacy, as it operates against the very principles that underpin privacy protection. Underscoring the importance of data scraping to AI and innovation more broadly, the authors suggest reconciling data scraping with data privacy by rethinking and reforming extant privacy laws.

Summary

Main Finding

Data scraping is central to AI model training but conflicts with core data‑privacy principles. Under U.S. law, scraping publicly accessible information is generally permissible; under EU law, scraping publicly available personal data will usually breach the GDPR unless a lawful basis under Article 6 applies. Although scraping is not categorically illegal in the EU, GDPR’s strict requirements make most large‑scale scraping of personal data difficult to justify in practice. Enforcement across EU member states is uneven. The authors conclude that existing privacy regimes inadequately balance AI innovation and individual privacy and call for legal reform to reconcile these objectives.

Key Points

  • Importance of data: AI model performance heavily depends on large, diverse, and (ideally) high‑quality datasets; this drives large‑scale automated scraping.
  • Forms of training data: public web data, licensed datasets, user‑provided data, third‑party contractual data, and data donations/altruism.
  • Nature of scraping: automated scraping (bots, crawlers) enables massive, low‑cost extraction distinct from manual copying; often occurs without data‑subject or platform consent.
  • U.S. legal position: courts and practice tend to allow scraping of publicly accessible information; absence of comprehensive federal privacy law creates permissive environment.
  • EU legal position: GDPR applies to personal data even if publicly available; lawful processing requires an Article 6 basis (consent, contract, legal obligation, vital interests, public task, or legitimate interests with balancing). Data scraping of personal data will typically fail GDPR tests unless designed to fit a lawful basis and meet principles (transparency, purpose limitation, minimization, security).
  • Privacy principles undermined by scraping: fairness; control and individual rights (access, correction, deletion); transparency and consent; purpose specification and restrictions on secondary use; data minimization; onward transfer/contractual protections; data security.
  • Ripple effects and harms: enables mass and shadow profiling, targeted cyberattacks, identity fraud, mass surveillance, and algorithmic/personalized price discrimination.
  • Practical caveats: GDPR enforcement varies across member states, producing heterogeneous compliance burdens and regulatory risk.
  • Policy recommendation (authors): rethink and reform privacy laws to better reconcile the needs of AI development with privacy protection.

Data & Methods

  • Methodology: doctrinal legal research (legal and jurisprudential analysis) and literature review of academic work, policy reports, and case law.
  • Legal sources and cases referenced: GDPR (esp. Article 6 and Article 5 principles); EU Data Governance Act (data altruism); U.S. cases and decisions cited historically include Authors Guild v. Google Inc. (2015) and Authors Guild v. HathiTrust (2014) (copyright contexts that informed scraping debates); Clearview AI litigation and regulatory scrutiny; Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson v. Anthropic (2025) referenced as an AI-related dispute over data sourcing; regulatory guidance from DPAs and reports from OECD, national authorities, and privacy scholars (Solove & Hartzog, etc.).
  • Empirical content: mainly legal interpretation and normative argumentation rather than econometric or experimental data.

Implications for AI Economics

  • Data access and cost effects:
    • Regulatory differences (U.S. permissive vs. EU restrictive) create cross‑jurisdictional asymmetries in effective access to training data.
    • Stricter EU requirements increase firms’ compliance costs (legal review, consent mechanisms, data audits, DPIAs) and transaction costs for acquiring lawful datasets (licenses, contracts).
    • Higher compliance costs favor incumbents with legal, financial, and data assets — increasing barriers to entry and potential market concentration in AI.
  • Innovation vs. privacy trade‑off:
    • Blanket bans or excessive restriction on scraping can slow model improvement, especially for smaller firms and researchers who rely on publicly available web data.
    • Conversely, permissive regimes risk negative externalities (privacy harms, fines, reputational risk) that could reduce social welfare and trust, ultimately undermining AI adoption.
  • Market design and data markets:
    • Pressure for regulated data access models: standardized licensing, APIs with clear terms, compensated data markets, data trusts, or safe‑harbors for research use could lower transaction costs while embedding privacy safeguards.
    • Data altruism / donation frameworks (e.g., EU Data Governance Act) may supply some lawful data but scale and representativeness are uncertain — affecting model performance and bias.
  • Algorithmic harms and economic welfare:
    • Scraped data enabling profiling and personalized pricing can produce redistributive harms and welfare losses (consumer surplus extraction, discrimination), prompting potential regulatory intervention that affects pricing strategies and market structure.
  • Compliance uncertainty and investment:
    • Heterogeneous enforcement (within EU and between US/EU) creates regulatory uncertainty that can deter investment in cross‑border model training and push firms to localize datasets or tailor models per jurisdiction.
  • Practical mitigation strategies with economic consequences:
    • Technical measures (differential privacy, synthetic data, federated learning) can reduce privacy risks but add development costs and may reduce model accuracy — an economic trade‑off between privacy protection and performance.
    • Licensing and negotiated data sharing shift costs to buyers and may create new revenue streams for data holders (platforms, publishers), potentially changing incumbents’ incentives.
  • Policy implications for regulators and policymakers (economically salient):
    • Consider harmonized, proportionate legal clarifications (e.g., narrow, supervised research exceptions; standardized lawful bases for AI training) to reduce fragmentation and lower compliance costs while protecting privacy.
    • Promote market mechanisms (data marketplaces, standardized APIs, certified processors) and technical standards (privacy‑preserving ML) to internalize privacy externalities without killing beneficial innovation.
    • Monitor distributional impacts: support smaller firms/researchers (grants, access programs) to avoid undue concentration of AI capabilities.

Concise takeaway: Large‑scale scraping materially lowers the marginal cost of collecting training data, accelerating AI development but producing significant privacy externalities and regulatory risk. Absent careful reform (legal, market, and technical), the economics of AI will favor established firms and create social harms (profiling, discrimination), so policy design should aim to align data access incentives with robust privacy safeguards.

Assessment

Paper Typedescriptive Evidence Strengthn/a — The article is a doctrinal legal review and comparative analysis, not an empirical study; it does not provide causal or quantitative evidence about economic outcomes. Methods Rigormedium — The paper applies standard doctrinal methods—statutory interpretation, case law review, and synthesis of secondary literature—systematically comparing U.S. and EU frameworks; however, it lacks empirical validation, primary data collection, or formal legal-theoretical innovation. SampleDoctrinal/legal sources and secondary literature: EU law (GDPR, Article 6), U.S. federal and state privacy statutes (discussed generally, including references to CCPA, BIPA), case law (e.g., Clearview AI litigation, Authors Guild v. Google, Anthropic litigation referenced), regulator guidance and DPAs' statements, and academic/policy literature on scraping and AI training. Themesgovernance innovation GeneralizabilityLimited to legal interpretation in U.S. and EU contexts; other jurisdictions (e.g., China, India, Africa, Latin America) are not covered., Findings reflect law and enforcement regimes as of the paper's writing; rapidly evolving case law and legislation may change conclusions., Does not provide empirical estimates of economic impacts (productivity, wages, firm performance), so implications for economic outcomes are inferential., Practical enforcement heterogeneity across EU member states limits uniform application of conclusions even within the EU.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In the United States, scraping publicly accessible data for AI training may be legally permissible. Governance And Regulation positive Legal permissibility of scraping publicly accessible data
Reading fidelity high
Study strength medium
not reported
0.18
Under European Union data protection law, scraping publicly available personal data violates data protection requirements unless the scraping satisfies a lawful basis under Article 6 of the GDPR. Regulatory Compliance negative Compliance of AI-related scraping with EU data protection law
Reading fidelity high
Study strength medium
not reported
0.18
Although data scraping is not inherently illegal under EU law, the authors conclude that it is almost unjustifiable in practice because of the stringent requirements of Article 6 of the GDPR. Governance And Regulation negative Practical justifiability of data scraping under EU law
Reading fidelity high
Study strength medium
not reported
0.18
Enforcement of GDPR provisions concerning data scraping has not been uniform across EU member states, with some jurisdictions taking a comparatively lenient approach. Governance And Regulation mixed Consistency and stringency of GDPR enforcement
Reading fidelity high
Study strength low
not reported
0.09
The authors argue that data scraping generally undermines data privacy because it conflicts with fundamental privacy-protection principles. Ai Safety And Ethics negative Protection of personal privacy
Reading fidelity high
Study strength low
not reported
0.09
Making personal information publicly available does not automatically constitute consent to its subsequent scraping, dissemination, or use for other purposes. Ai Safety And Ethics negative Validity of consent and privacy expectations for publicly available personal data
Reading fidelity high
Study strength medium
not reported
0.18
Data scraping can enable AI-driven profiling and increase the capacity for mass profiling and mass surveillance. Ai Safety And Ethics negative Risk of AI-driven profiling and mass surveillance
Reading fidelity high
Study strength low
not reported
0.09
Data scraping can contribute to personalized or algorithmic price discrimination by enabling businesses to collect and analyze detailed individual-level data to build consumer profiles. Consumer Welfare negative Risk of differential pricing based on personal characteristics
Reading fidelity high
Study strength low
not reported
0.09
The quality and quantity of training data affect AI model performance, making data collection a prerequisite for effective AI model development. Output Quality positive AI model performance and training effectiveness
Reading fidelity high
Study strength low
not reported
0.09

Notes