The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

LLM tokenizers systematically over-segment Arabic, inflating token usage by up to fourfold and mechanically raising inference costs and constraining context; this infrastructural bias—rooted in tokenization, attention, and optimization—creates unequal access that is not resolved by dataset fixes alone.

The Algorithmic Unconscious: Structural Mechanisms and Implicit Biases in Large Language Models
Philippe Boisnard · February 08, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Philippe Boisnard unresolved corpus identity

Semantic Scholar

Latest observation:

  1. P. Boisnard provider ID
Tokenization regimes in multiple LLM infrastructures systematically over-segment Arabic (MSA and Maghrebi), inflating token counts by roughly 1.6x–4x versus English and thereby mechanically raising inference costs, shrinking effective context capacity, and producing an infrastructural form of bias independent of dataset composition.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This article introduces the concept of the algorithmic unconscious to designate the set of structural determinations that operate within large language models (LLMs) without being accessible either to the model's own reflexivity or to that of its users. In contrast to approaches that reduce AI bias solely to dataset composition or to the projection of human intentionality, we argue that a significant class of biases emerges directly from the technical mechanisms of the models themselves: tokenization, attention, statistical optimization, and alignment procedures. By framing bias as an infrastructural phenomenon, this approach resolves a central theoretical ambiguity surrounding responsibility, neutrality, and correction in contemporary LLMs. Based on a comparative analysis of tokenization across a corpus of parallel sentences, we show that Arabic languages (Modern Standard Arabic and Maghrebi dialects) undergo a systematic inflation in token count relative to English, with ratios ranging from 1.6x to nearly 4x depending on the infrastructure (OpenAI, Anthropic, SentencePiece/Mistral). This over-segmentation constitutes a measurable infrastructural bias that mechanically increases inference costs, constrains access to contextual space, and alters attentional weighting within model representations. We relate these empirical findings to three additional structural mechanisms: causal bias (correlation vs causation), the erasure of minoritized features through dimensional collapse, and normative biases induced by safety alignment. Finally, we propose a framework for a technical clinic of models, grounded in the audit of tokenization regimes, latent space topology, and alignment systems, as a necessary condition for the critical appropriation of AI infrastructures.

Summary

Main Finding

Philippe Boisnard introduces the concept of the "algorithmic unconscious": structural, machine-internal mechanisms in large language models (LLMs) — notably tokenization, attention, statistical optimization, and alignment procedures — that generate systematic biases independent of dataset composition or human intent. Empirically, tokenization produces measurable infrastructural bias: Arabic varieties (Modern Standard Arabic and Maghrebi dialects) are systematically over-segmented relative to English (token inflation ≈ 1.6× for some OpenAI pipelines up to ≈ 4× for SentencePiece/Mistral), producing higher inference/training costs, reduced effective context, and distorted attentional and representational geometry. Boisnard links this token-level effect to three further structural mechanisms — causal bias (correlation ≠ causation), dimensional collapse (erasure of minoritized features), and safety-alignment biases — arguing these jointly reproduce and amplify linguistic and cultural hierarchies.

Key Points

  • Definition: Algorithmic unconscious = structural, non-reflexive elements embedded in model architectures and pipelines (tokenizers, attention, optimization, alignment) that produce biases invisible to both users and models themselves.
  • Tokenization bias (infrastructural bias):
    • Comparative tokenization of parallel sentences shows Arabic (incl. Darija) yields 1.6×–4× more tokens than English depending on tokenizer/model infrastructure.
    • Finer segmentation for under-resourced languages increases token fertility, destabilizes representations, and increases per-query costs.
    • Different tokenization logics across models (e.g., SentencePiece vs. BPE) produce non-uniform biases; Mistral (SentencePiece) often yields finer segmentation than ChatGPT-style BPE.
  • Attention & contextual weighting:
    • Token frequency, morphological granularity, and subword choices shape internal “geographies” of attention where majority languages become central attractors and minority tokens peripheral or noisy.
  • Causal bias:
    • LLMs learn conditional distributions, not interventions—hence they can reinterpret correlations as causal relations, leading to spurious or stereotyped associations, particularly for underrepresented cultures whose contexts are sparse.
  • Dimensional collapse (erasure):
    • Optimization pressures favor frequent tokens/contexts; minority-language embeddings occupy reduced, anisotropic subspaces, losing dialectal and cultural specificity (e.g., Darija collapsing toward MSA).
    • This is analogous to mode collapse: diversity is lost in favor of statistically "safe" attractors.
  • Safety-alignment biases:
    • Filtering, RLHF/RLAIF, and moderation pipelines embed normative choices (often Western/North American) that have asymmetric effects across languages, compounding representational harms.
  • Social/economic consequences:
    • Higher token use → higher training/inference costs → access and affordability gaps for speakers of marginalized languages.
    • Epistemic and cultural erasure: LLM outputs can normalize majority frameworks, producing computational colonialism.
  • Proposal:
    • A "technical clinic" for models is required, centered on audits of tokenization, latent-space topology, and alignment systems to make these structural biases visible and actionable.

Data & Methods

  • Empirical component:
    • Comparative tokenization study on a corpus of parallel sentences across languages (English, Modern Standard Arabic, Maghrebi dialects/Darija).
    • Token counts measured across multiple infrastructures/tokenizers (OpenAI pipelines, Anthropic, SentencePiece/Mistral).
    • Reported token-inflation ratios: ~1.6× (OpenAI) up to nearly 4× (SentencePiece/Mistral) for Arabic vs. English on the sample corpus.
    • Qualitative/quantitative inspection of segmentation granularity (appendices: character-level and subword-level examples).
  • Analytical methods:
    • Embedding/latent-space inspection and literature synthesis on anisotropy, dimensional collapse, and representational geometry (citing Ethayarajh, Jing et al., Naous & Xu).
    • Conceptual framing of attention weight distributions and optimization dynamics to explain emergent erasure and causal reductionism.
    • Review of alignment/moderation pipelines and impact analyses (citing recent studies on censorship, moderation bias, and RLHF).
  • Sources and triangulation:
    • Uses prior empirical studies (Bari et al., Alyafeai et al., Teklehaymanot & Nejdl, Petrov et al.) alongside new tokenization counts and comparative tokenizer analysis.
  • Limitations acknowledged:
    • Proprietary and evolving tokenizers/model versions introduce variability; token ratios depend on tokenizer version, vocabulary, and pre-/post-processing.
    • Corpus size, representativeness of parallel sentences, and heterogeneity within dialects (Darija variants) constrain generalizability.
    • Some claims (e.g., exact downstream attention effects) are argued conceptually and by proxy from latent-space phenomena rather than exhaustive causal experiments.

Implications for AI Economics

  • Direct cost implications:
    • Token inflation for under-resourced languages multiplies both training and inference compute and monetary costs (Boisnard cites up to ~4× inflation in token usage). This reduces cost-efficiency and raises per-query prices for speakers of those languages under usage-based pricing models.
    • Increased compute for minority-language support creates negative incentives for platform providers to optimize for those languages, reinforcing market concentration around high-resource languages.
  • Market and access effects:
    • Higher operational costs and lower quality for minority languages deepen digital divides: decreased affordability, lower service quality, and lower adoption in research, education, and local markets.
    • Economies of scale favor models tuned to majority languages; public goods or subsidized models may be needed to correct market failures.
  • Value capture and extraction:
    • "Computational colonialism": models structurally valorize data and forms of expression aligned with dominant cultures, extracting value (attention, user activity, content) while reducing representational and cultural returns to marginalized communities.
    • Businesses serving minority-language communities face higher costs to achieve parity, creating barriers to competition and localization.
  • Product design and pricing:
    • Usage-based billing (tokens) can be regressive across language groups: identical semantic queries cost more in tokens for certain languages. This may be considered discriminatory pricing externality embedded in infrastructure.
    • Tokenization-aware pricing, language-aware quotas, or subsidized contexts may be necessary short-term interventions.
  • Policy, regulation, and audit economics:
    • Transparency and auditing requirements (tokenization audits, latent-space topology reviews, alignment impact assessments) create compliance costs but are essential to correct structural bias.
    • Regulators may require model providers to report language-wise tokenization efficiency, representational metrics, and differential moderation impacts — shifting costs into compliance and operational reporting.
  • R&D and investment implications:
    • There's economic opportunity in developing multilingual tokenizers, language-specific models, or adapter layers that reduce token fertility and preserve dialectal features — a new market for localization tech and specialist LLMs.
    • Funding (public or philanthropic) may be justified to support modelling and corpora creation for under-resourced languages to correct market under-provisioning.
  • Long-run epistemic economy:
    • Because LLMs are increasingly integrated into research and education, structural biases can reshape which knowledge is producible and economically valuable (e.g., research outputs, localized content production), with distributional effects on cultural industries and intellectual labor markets.

Suggestions implied by the paper for economists and policymakers: - Measure and disclose token-fertility and representation metrics by language as part of model audits. - Consider subsidies, public models, or regulation to correct for market failures that leave under-resourced languages under-supported. - Re-evaluate usage- and token-based pricing models that embed infrastructural bias and create unequal access. - Invest in tooling and standards for auditing tokenizers, alignment pipelines, and latent-space geometry to quantify representational and economic harms.

— End of summary.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides clear, directly measured differences in token counts across parallel corpora and multiple tokenization infrastructures, yielding robust descriptive evidence of a measurable phenomenon; however it does not link these differences to quantified downstream economic outcomes (e.g., actual inference cost increases at scale, user access metrics, or effects on model performance), and potential confounders (corpus domain, tokenizer versions, sample size) are not fully addressed. Methods Rigormedium — The core empirical method—comparative analysis of tokenization on a parallel sentence corpus across multiple tokenizers—is an appropriate and replicable descriptive approach; rigor is limited by likely unspecified sample size and representativeness, lack of sensitivity analyses (e.g., varying corpora, tokenizer/hyperparameter versions), and absence of direct downstream validation (cost/performance impacts, user studies). SampleA corpus of parallel sentences in English, Modern Standard Arabic, and Maghrebi Arabic dialects, tokenized with multiple infrastructures (proprietary tokenizers tied to OpenAI and Anthropic, and open-source SentencePiece as used by Mistral); analyses report token count ratios ranging ~1.6x–4x (Arabic vs English) depending on tokenizer. Themesinequality adoption productivity GeneralizabilityResults based on specific languages (Modern Standard Arabic and Maghrebi dialects) — may not generalize to other non-Latin scripts or languages with different morphological properties., Tokenization behaviour depends on tokenizer versions and hyperparameters; findings may change with updated tokenizers or model-specific pretokenization pipelines., Corpus domain and register (type of parallel sentences) may affect tokenization ratios; results may not hold for other text genres (social media, code, speech transcripts)., Paper documents infrastructural bias but does not empirically measure downstream economic impacts (inference cost, latency, access), limiting direct inference to economic outcomes., Proprietary tokenizers and closed model internals limit replicability and comparison across all production systems.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The paper introduces the concept of the "algorithmic unconscious" to designate the set of structural determinations that operate within large language models (LLMs) without being accessible either to the model's own reflexivity or to that of its users. Ai Safety And Ethics negative existence of inaccessible structural determinations within LLMs
Reading fidelity high
Study strength speculative
not reported
0.03
A significant class of biases emerges directly from the technical mechanisms of models themselves—specifically tokenization, attention, statistical optimization, and alignment procedures—rather than solely from dataset composition or the projection of human intentionality. Ai Safety And Ethics negative source/origin of bias in LLMs (technical mechanisms vs dataset/human projection)
Reading fidelity high
Study strength medium
not reported
0.18
Framing bias as an infrastructural phenomenon resolves a central theoretical ambiguity surrounding responsibility, neutrality, and correction in contemporary LLMs. Governance And Regulation positive clarity in responsibility/neutrality/correction frameworks for LLMs
Reading fidelity high
Study strength speculative
not reported
0.03
A comparative analysis of tokenization across a corpus of parallel sentences shows that Arabic languages (Modern Standard Arabic and Maghrebi dialects) undergo a systematic inflation in token count relative to English, with ratios ranging from 1.6x to nearly 4x depending on the infrastructure (OpenAI, Anthropic, SentencePiece/Mistral). Task Allocation negative token count inflation for Arabic relative to English
Reading fidelity high
Study strength medium
1.6x to nearly 4x
0.18
This over-segmentation (inflated token counts for Arabic) constitutes a measurable infrastructural bias that mechanically increases inference costs, constrains access to contextual space, and alters attentional weighting within model representations. Organizational Efficiency negative inference costs; access to contextual space; changes in attentional weighting
Reading fidelity high
Study strength medium
not reported
0.18
Three additional structural mechanisms contribute to infrastructural bias: causal bias (correlation vs causation), the erasure of minoritized features through dimensional collapse, and normative biases induced by safety alignment. Ai Safety And Ethics negative mechanisms producing bias (causal misinterpretation, feature erasure, alignment norms)
Reading fidelity high
Study strength medium
not reported
0.18
The authors propose a 'technical clinic' framework for auditing models—grounded in the audit of tokenization regimes, latent space topology, and alignment systems—as a necessary condition for the critical appropriation of AI infrastructures. Governance And Regulation positive feasibility/effectiveness of auditing frameworks for AI infrastructures
Reading fidelity high
Study strength speculative
not reported
0.03

Notes