0 cumulative citations
View corpus contextGenerative AI hardwires majority ways of knowing into the architecture of knowledge, structurally marginalising minority cultures; current anti-discrimination, cultural-rights and media laws, which target outputs, fail to reach this upstream harm, so regulation must shift to govern training data, alignment methods and model architecture.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expression into a single probabilistic model whose statistical baseline reflects the languages, assumptions, and cultural frameworks of the dominant culture. Minority epistemologies are not excluded but absorbed: present in the training data, yet structurally subordinated in the output. The result is not a collection of discrete biases that can be audited and corrected. It is an epistemic condition embedded in the architecture from which all outputs emerge. This unified harm cuts across three legal domains -- anti-discrimination law, cultural and linguistic rights, and democratic viewpoint pluralism -- and each fails to address it for the same structural reason: existing law regulates downstream, at the level of decisions and applications. The remedy must match the site of harm. If epistemic subordination is produced at the level of model training, then law must learn to govern at that level.
Summary
Main Finding
Generative AI creates a structural epistemic bias — “epistemic subordination” — by encoding majority cultural and linguistic ways of knowing into the foundational parameters of large models. Because training corpora are internet-dominant and model internals are effectively irreversible, downstream fixes (content filters, RLHF) can change outputs but not the model’s epistemic baseline. This produces harms that cut across discrimination, minority cultural/linguistic/religious rights, and democratic viewpoint pluralism — and these harms are not well addressed by existing downstream legal or regulatory instruments. Law and policy must therefore shift upstream to the infrastructure of model creation (data composition, training methods, architecture).
Key Points
- Definition: Epistemic subordination — the architectural encoding of the majority’s epistemology as the statistical default of a model, with other epistemologies treated as deviations.
- Three-stage mechanism:
- Training data: Foundation models are trained on large internet corpora that overrepresent Western, English-language, and digitally privileged voices.
- Irreversibility/opacity: Training compresses data into billions of parameters; the contribution of specific sources cannot be isolated or undone.
- Downstream alignment limits: Value-alignment (RLHF, system prompts, content moderation) changes outputs but does not reconstruct the underlying epistemic representation.
- Legal analysis:
- Anti-discrimination law: Doctrine (intent, disparate impact) presumes identifiable practices/criteria; generative AI’s architectural bias cannot be targeted by existing disparate impact frameworks because the “criterion” is the entire model.
- Cultural/linguistic/religious rights: These rights protect institutional spaces where minorities can reproduce their epistemic practices. Models embed bias in infrastructure, which minority communities cannot realistically recreate (high training cost, data scarcity), undermining those protections.
- Viewpoint pluralism/democracy: Public deliberation depends on exposure to diverse epistemic frameworks. If AI is a dominant knowledge infrastructure, it homogenizes the epistemic baseline and constrains public discourse in ways current free-speech or media regulation cannot remedy.
- Empirical examples cited: LLMs associating African American English with negative traits; medical LLMs exhibiting sociodemographic biases; text-to-image models amplifying racial/professional stereotypes; multilingual performance gaps (e.g., strong English accuracy vs. low accuracy in many African languages); failure modes when alignment attempts to force “diversity” (e.g., absurd image generations).
- Normative claim: Remedies must target model construction (data composition, training architecture, alignment methods) rather than only outputs or deployment.
Data & Methods
- Nature of the paper: conceptual and legal-theoretical essay synthesizing technical characteristics of generative models, empirical studies, and doctrinal analysis. Not an original empirical study; uses interdisciplinary literature.
- Technical evidence and references used:
- Training corpora descriptions (e.g., GPT-3’s filtered Common Crawl and book/Wikipedia corpora).
- Technical properties of deep learning/opacity (Chesterman on opacity; deep learning as nontransparent statistical compression).
- Alignment techniques: RLHF, constitutional AI, system-level instructions and content moderation pipelines.
- Empirical studies cited (examples):
- Hofmann et al., Nature (2024) — dialect-based stereotyping in LLM outputs.
- Omar et al., Nature Medicine (2025) — sociodemographic biases in medical LLM decision-making.
- Bianchi et al., FAccT (2023) — demographic stereotyping in text-to-image models.
- MMLU-ProX and multilingual benchmarking (2025) — cross-linguistic performance gaps; larger accuracy in English vs. many African languages.
- Case examples of alignment failures (journalistic/academic reports on image generation failures).
- Legal method: doctrinal mapping showing why existing anti-discrimination, cultural-rights, and viewpoint-pluralism doctrines regulate downstream conduct and outputs, not upstream architecture; thus they structurally fail to capture epistemic-subordination harms.
- Limitations: The essay is primarily conceptual and normative, grounded in cited empirical work but not offering new large-scale empirical quantification of epistemic subordination across markets or countries.
Implications for AI Economics
- Market structure & public-goods framing:
- Foundation models are quasi-public goods / infrastructures: high fixed (training) costs, concentrated suppliers, nonrival use, strong network effects (standardization of epistemic baseline).
- Epistemic subordination is an externality: the informational and cultural harms to minorities and democratic debate are not internalized by model developers.
- Barrier to entry: massive data and compute needs make it infeasible for minority communities or smaller firms to create alternative epistemic infrastructure, reinforcing incumbents’ market power.
- Distributional and welfare effects:
- Consumer surplus is uneven: dominant groups receive higher-quality epistemic services; minorities receive lower-quality/biased information, reducing their informational welfare, human capital accumulation, and economic opportunities.
- Cultural capital depreciation: minority languages and knowledge systems face erosion, with potential long-term negative effects on cultural industries, education outcomes, and labor-market participation.
- Political economy risk: homogenized epistemic baselines may reduce policy diversity, crowd out minority policy narratives, and change electoral information environments — potential negative externalities on social welfare not captured in typical GDP metrics.
- Measurement and empirical research agenda for economists:
- Develop metrics of “epistemic diversity” and model-aligned bias (e.g., cross-lingual performance gaps, alignment of generated content with majority vs. minority norms).
- Quantify market concentration and entry costs in foundation-model markets; measure correlation between model-market share and cross-group welfare gaps.
- Estimate welfare loss from epistemic subordination (lost cultural production, reduced trust, information-quality harms).
- Design experimental and quasi-experimental studies to identify causal effects of model usage on language vitality, political preferences, or labor outcomes.
- Policy and regulatory instruments (economics-relevant):
- Upstream data mandates and composition rules: require disclosure of training-data composition or minimum representation thresholds for protected languages/epistemologies; may be implemented via audits or standardized benchmarks.
- Subsidies and public funding: subsidize the creation of minority-language corpora and culturally-specific datasets; fund public-interest foundation models trained with inclusive data.
- Procurement and public provision: governments purchase or build multilingual, pluralistic public models to supply education, translation, and public services — correcting market undersupply.
- Data trusts and stewardship: create community data trusts that aggregate and license minority cultural data under governance that captures collective preferences and economic returns.
- Antitrust and competition policy: scrutinize vertical integration between model builders and downstream platforms; facilitate model interoperability and data portability to lower entry barriers.
- Standards and certification: create epistemic-diversity benchmarks and certification regimes that buyers and regulators can use; tie procurement or tax advantages to certified models.
- R&D incentives: tax credits or prizes for models that demonstrably reduce epistemic-subordination metrics (e.g., parity across languages/communities).
- Liability and disclosure: require provenance metadata, model-performance disclosures across demographic and linguistic groups, and penalties for demonstrable harm where feasible.
- Trade-offs and design considerations:
- Cost vs. inclusivity: mandating diverse training data raises compliance costs and can favor incumbents unless paired with subsidies/public provision.
- Measurement difficulty: attributing harms to model architecture vs. downstream deployment is empirically challenging; regulation may need to rely on proxies and benchmarks.
- Innovation vs. regulation balance: heavy upstream mandates could slow innovation; targeted public funding and incentives can mitigate distributional harms without unduly constraining experimentation.
- Practical near-term steps for policymakers and economists:
- Fund corpora-building efforts for underrepresented languages and communities; integrate these corpora into public models and open benchmarks.
- Establish standard multilingual/epistemic-equality evaluation suites for procurement and certification.
- Pilot public foundation models focused on minority needs (education, legal aid, cultural preservation) and measure impact.
- Incorporate epistemic-diversity externalities into cost–benefit analyses of AI regulation and competition policy.
- Support interdisciplinary research combining law, economics, and NLP to operationalize metrics of epistemic subordination and to evaluate policy interventions.
In short: epistemic subordination reframes foundation models as infrastructure with social externalities that typical downstream regulation and market forces do not correct. From an AI-economics perspective, remedies should combine upstream regulatory requirements, public provision/subsidies, competition policy, and new measurement standards to correct market failures, preserve plural epistemic infrastructures, and protect welfare across groups.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Generative AI encodes the majority's way of knowing as the infrastructure of knowledge itself, subordinating other epistemologies to a statistical default. Ai Safety And Ethics | negative | Relative representation and status of majority and minority epistemologies in generative-AI knowledge systems |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Internet-based training data systematically overrepresents English-language, Western, and digitally privileged populations, causing generative-AI models to encode majority linguistic and cultural frameworks as their statistical default. Ai Safety And Ethics | negative | Cultural and linguistic representation in model training data and model defaults |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The statistical compression of training data into billions of model parameters makes the model's biased baseline effectively irreversible and prevents outputs from being practically traced to identifiable training sources. Ai Safety And Ethics | negative | Traceability and inspectability of model knowledge and bias |
Reading fidelity
high
Study strength
medium
|
not reported
|
| RLHF, constitutional AI, system instructions, and content moderation primarily change model outputs and do not restructure the epistemic foundation established during pretraining. Ai Safety And Ethics | mixed | Effect of alignment methods on model knowledge and epistemic foundations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Generative AI systems have produced historically absurd outputs when instructed to generate culturally diverse images, including racially diverse Nazi soldiers. Output Quality | negative | Historical and contextual accuracy of generated images |
Reading fidelity
high
Study strength
low
|
not reported
|
| Large language models associate African American English with negative stereotypes, generating descriptors such as 'dirty,' 'lazy,' and 'aggressive' in response to the dialect itself. Ai Safety And Ethics | negative | Stereotypical sentiment and descriptors generated for different dialects |
Reading fidelity
high
Study strength
high
|
not reported
|
| The medical AI systems discussed in the essay evaluate LGBTQIA+ patients for mental-health concerns at six to seven times the clinical baseline and recommend less aggressive diagnostic imaging for lower-income patients. Decision Quality | negative | Mental-health evaluation rates and recommended diagnostic-imaging intensity across demographic and income groups |
Reading fidelity
high
Study strength
medium
|
six to seven times the clinical baseline
|
| Text-to-image generators amplify demographic stereotypes, associating high-status professions with lighter skin and criminality with darker skin. Ai Safety And Ethics | negative | Demographic associations in generated images for occupations and criminality |
Reading fidelity
high
Study strength
high
|
not reported
|
| The best-performing multilingual language models show substantial cross-linguistic accuracy disparities, with 80.7% accuracy in English versus 57.0% in Yoruba and 58.6% in Wolof. Output Quality | negative | Benchmark accuracy across languages |
Reading fidelity
high
Study strength
high
|
80.7% accuracy in English; 57.0% in Yoruba; 58.6% in Wolof; gap up to 24.3 percentage points
|
| Generative AI can introduce majority cultural and normative assumptions into minority institutional settings, such as secular-Western reasoning in religious-school curricula and marginalization of minority languages in classrooms. Ai Safety And Ethics | negative | Cultural and linguistic autonomy of minority institutions using AI-assisted knowledge tools |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| If generative AI becomes a central knowledge infrastructure, its culturally homogenized defaults may constrain the frameworks available for public debate and democratic deliberation. Governance And Regulation | negative | Diversity of frameworks and viewpoints available in democratic public discourse |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Existing anti-discrimination law, cultural and linguistic rights, viewpoint-pluralism doctrines, and AI-specific regulation generally operate downstream on outputs or applications and therefore do not adequately address epistemic subordination produced during model construction. Governance And Regulation | negative | Capacity of existing legal frameworks to address structural epistemic harms in generative-AI development |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| An adequate legal response to epistemic subordination would need to regulate the model-construction process, including training-data composition, alignment methods, and model architecture. Governance And Regulation | positive | Regulatory targeting of upstream generative-AI development processes |
Reading fidelity
high
Study strength
speculative
|
not reported
|