0 cumulative citations
View corpus contextLLM tokenizers systematically over-segment Arabic, inflating token usage by up to fourfold and mechanically raising inference costs and constraining context; this infrastructural bias—rooted in tokenization, attention, and optimization—creates unequal access that is not resolved by dataset fixes alone.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This article introduces the concept of the algorithmic unconscious to designate the set of structural determinations that operate within large language models (LLMs) without being accessible either to the model's own reflexivity or to that of its users. In contrast to approaches that reduce AI bias solely to dataset composition or to the projection of human intentionality, we argue that a significant class of biases emerges directly from the technical mechanisms of the models themselves: tokenization, attention, statistical optimization, and alignment procedures. By framing bias as an infrastructural phenomenon, this approach resolves a central theoretical ambiguity surrounding responsibility, neutrality, and correction in contemporary LLMs. Based on a comparative analysis of tokenization across a corpus of parallel sentences, we show that Arabic languages (Modern Standard Arabic and Maghrebi dialects) undergo a systematic inflation in token count relative to English, with ratios ranging from 1.6x to nearly 4x depending on the infrastructure (OpenAI, Anthropic, SentencePiece/Mistral). This over-segmentation constitutes a measurable infrastructural bias that mechanically increases inference costs, constrains access to contextual space, and alters attentional weighting within model representations. We relate these empirical findings to three additional structural mechanisms: causal bias (correlation vs causation), the erasure of minoritized features through dimensional collapse, and normative biases induced by safety alignment. Finally, we propose a framework for a technical clinic of models, grounded in the audit of tokenization regimes, latent space topology, and alignment systems, as a necessary condition for the critical appropriation of AI infrastructures.
Summary
Main Finding
Philippe Boisnard introduces the concept of the "algorithmic unconscious": structural, machine-internal mechanisms in large language models (LLMs) — notably tokenization, attention, statistical optimization, and alignment procedures — that generate systematic biases independent of dataset composition or human intent. Empirically, tokenization produces measurable infrastructural bias: Arabic varieties (Modern Standard Arabic and Maghrebi dialects) are systematically over-segmented relative to English (token inflation ≈ 1.6× for some OpenAI pipelines up to ≈ 4× for SentencePiece/Mistral), producing higher inference/training costs, reduced effective context, and distorted attentional and representational geometry. Boisnard links this token-level effect to three further structural mechanisms — causal bias (correlation ≠ causation), dimensional collapse (erasure of minoritized features), and safety-alignment biases — arguing these jointly reproduce and amplify linguistic and cultural hierarchies.
Key Points
- Definition: Algorithmic unconscious = structural, non-reflexive elements embedded in model architectures and pipelines (tokenizers, attention, optimization, alignment) that produce biases invisible to both users and models themselves.
- Tokenization bias (infrastructural bias):
- Comparative tokenization of parallel sentences shows Arabic (incl. Darija) yields 1.6×–4× more tokens than English depending on tokenizer/model infrastructure.
- Finer segmentation for under-resourced languages increases token fertility, destabilizes representations, and increases per-query costs.
- Different tokenization logics across models (e.g., SentencePiece vs. BPE) produce non-uniform biases; Mistral (SentencePiece) often yields finer segmentation than ChatGPT-style BPE.
- Attention & contextual weighting:
- Token frequency, morphological granularity, and subword choices shape internal “geographies” of attention where majority languages become central attractors and minority tokens peripheral or noisy.
- Causal bias:
- LLMs learn conditional distributions, not interventions—hence they can reinterpret correlations as causal relations, leading to spurious or stereotyped associations, particularly for underrepresented cultures whose contexts are sparse.
- Dimensional collapse (erasure):
- Optimization pressures favor frequent tokens/contexts; minority-language embeddings occupy reduced, anisotropic subspaces, losing dialectal and cultural specificity (e.g., Darija collapsing toward MSA).
- This is analogous to mode collapse: diversity is lost in favor of statistically "safe" attractors.
- Safety-alignment biases:
- Filtering, RLHF/RLAIF, and moderation pipelines embed normative choices (often Western/North American) that have asymmetric effects across languages, compounding representational harms.
- Social/economic consequences:
- Higher token use → higher training/inference costs → access and affordability gaps for speakers of marginalized languages.
- Epistemic and cultural erasure: LLM outputs can normalize majority frameworks, producing computational colonialism.
- Proposal:
- A "technical clinic" for models is required, centered on audits of tokenization, latent-space topology, and alignment systems to make these structural biases visible and actionable.
Data & Methods
- Empirical component:
- Comparative tokenization study on a corpus of parallel sentences across languages (English, Modern Standard Arabic, Maghrebi dialects/Darija).
- Token counts measured across multiple infrastructures/tokenizers (OpenAI pipelines, Anthropic, SentencePiece/Mistral).
- Reported token-inflation ratios: ~1.6× (OpenAI) up to nearly 4× (SentencePiece/Mistral) for Arabic vs. English on the sample corpus.
- Qualitative/quantitative inspection of segmentation granularity (appendices: character-level and subword-level examples).
- Analytical methods:
- Embedding/latent-space inspection and literature synthesis on anisotropy, dimensional collapse, and representational geometry (citing Ethayarajh, Jing et al., Naous & Xu).
- Conceptual framing of attention weight distributions and optimization dynamics to explain emergent erasure and causal reductionism.
- Review of alignment/moderation pipelines and impact analyses (citing recent studies on censorship, moderation bias, and RLHF).
- Sources and triangulation:
- Uses prior empirical studies (Bari et al., Alyafeai et al., Teklehaymanot & Nejdl, Petrov et al.) alongside new tokenization counts and comparative tokenizer analysis.
- Limitations acknowledged:
- Proprietary and evolving tokenizers/model versions introduce variability; token ratios depend on tokenizer version, vocabulary, and pre-/post-processing.
- Corpus size, representativeness of parallel sentences, and heterogeneity within dialects (Darija variants) constrain generalizability.
- Some claims (e.g., exact downstream attention effects) are argued conceptually and by proxy from latent-space phenomena rather than exhaustive causal experiments.
Implications for AI Economics
- Direct cost implications:
- Token inflation for under-resourced languages multiplies both training and inference compute and monetary costs (Boisnard cites up to ~4× inflation in token usage). This reduces cost-efficiency and raises per-query prices for speakers of those languages under usage-based pricing models.
- Increased compute for minority-language support creates negative incentives for platform providers to optimize for those languages, reinforcing market concentration around high-resource languages.
- Market and access effects:
- Higher operational costs and lower quality for minority languages deepen digital divides: decreased affordability, lower service quality, and lower adoption in research, education, and local markets.
- Economies of scale favor models tuned to majority languages; public goods or subsidized models may be needed to correct market failures.
- Value capture and extraction:
- "Computational colonialism": models structurally valorize data and forms of expression aligned with dominant cultures, extracting value (attention, user activity, content) while reducing representational and cultural returns to marginalized communities.
- Businesses serving minority-language communities face higher costs to achieve parity, creating barriers to competition and localization.
- Product design and pricing:
- Usage-based billing (tokens) can be regressive across language groups: identical semantic queries cost more in tokens for certain languages. This may be considered discriminatory pricing externality embedded in infrastructure.
- Tokenization-aware pricing, language-aware quotas, or subsidized contexts may be necessary short-term interventions.
- Policy, regulation, and audit economics:
- Transparency and auditing requirements (tokenization audits, latent-space topology reviews, alignment impact assessments) create compliance costs but are essential to correct structural bias.
- Regulators may require model providers to report language-wise tokenization efficiency, representational metrics, and differential moderation impacts — shifting costs into compliance and operational reporting.
- R&D and investment implications:
- There's economic opportunity in developing multilingual tokenizers, language-specific models, or adapter layers that reduce token fertility and preserve dialectal features — a new market for localization tech and specialist LLMs.
- Funding (public or philanthropic) may be justified to support modelling and corpora creation for under-resourced languages to correct market under-provisioning.
- Long-run epistemic economy:
- Because LLMs are increasingly integrated into research and education, structural biases can reshape which knowledge is producible and economically valuable (e.g., research outputs, localized content production), with distributional effects on cultural industries and intellectual labor markets.
Suggestions implied by the paper for economists and policymakers: - Measure and disclose token-fertility and representation metrics by language as part of model audits. - Consider subsidies, public models, or regulation to correct for market failures that leave under-resourced languages under-supported. - Re-evaluate usage- and token-based pricing models that embed infrastructural bias and create unequal access. - Invest in tooling and standards for auditing tokenizers, alignment pipelines, and latent-space geometry to quantify representational and economic harms.
— End of summary.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The paper introduces the concept of the "algorithmic unconscious" to designate the set of structural determinations that operate within large language models (LLMs) without being accessible either to the model's own reflexivity or to that of its users. Ai Safety And Ethics | negative | existence of inaccessible structural determinations within LLMs |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A significant class of biases emerges directly from the technical mechanisms of models themselves—specifically tokenization, attention, statistical optimization, and alignment procedures—rather than solely from dataset composition or the projection of human intentionality. Ai Safety And Ethics | negative | source/origin of bias in LLMs (technical mechanisms vs dataset/human projection) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Framing bias as an infrastructural phenomenon resolves a central theoretical ambiguity surrounding responsibility, neutrality, and correction in contemporary LLMs. Governance And Regulation | positive | clarity in responsibility/neutrality/correction frameworks for LLMs |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A comparative analysis of tokenization across a corpus of parallel sentences shows that Arabic languages (Modern Standard Arabic and Maghrebi dialects) undergo a systematic inflation in token count relative to English, with ratios ranging from 1.6x to nearly 4x depending on the infrastructure (OpenAI, Anthropic, SentencePiece/Mistral). Task Allocation | negative | token count inflation for Arabic relative to English |
Reading fidelity
high
Study strength
medium
|
1.6x to nearly 4x
|
| This over-segmentation (inflated token counts for Arabic) constitutes a measurable infrastructural bias that mechanically increases inference costs, constrains access to contextual space, and alters attentional weighting within model representations. Organizational Efficiency | negative | inference costs; access to contextual space; changes in attentional weighting |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Three additional structural mechanisms contribute to infrastructural bias: causal bias (correlation vs causation), the erasure of minoritized features through dimensional collapse, and normative biases induced by safety alignment. Ai Safety And Ethics | negative | mechanisms producing bias (causal misinterpretation, feature erasure, alignment norms) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The authors propose a 'technical clinic' framework for auditing models—grounded in the audit of tokenization regimes, latent space topology, and alignment systems—as a necessary condition for the critical appropriation of AI infrastructures. Governance And Regulation | positive | feasibility/effectiveness of auditing frameworks for AI infrastructures |
Reading fidelity
high
Study strength
speculative
|
not reported
|