The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models give brands different lives: Chinese LLMs mention brands in English queries about 30 percentage points more often than international models, creating an 'Existence Gap' that can make firms invisible to AI-mediated discovery.

Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
Huang Junyao, Situ Ruimin, Ye Renqin · December 30, 2025
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Huang Junyao unresolved corpus identity
  2. Situ Ruimin unresolved corpus identity
  3. Ye Renqin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Junyao Huang provider ID
  2. Situ Ruimin provider ID
  3. Renqin Ye provider ID
The paper documents large cross-model differences in brand visibility — Chinese LLMs mention brands in English queries far more often than International LLMs — and argues these gaps reflect training-data geography that creates market-entry barriers for underrepresented brands.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As artificial intelligence systems increasingly mediate consumer information discovery, brands face algorithmic invisibility. This study investigates Cultural Encoding in Large Language Models (LLMs) -- systematic differences in brand recommendations arising from training data composition. Analyzing 1,909 pure-English queries across 6 LLMs (GPT-4o, Claude, Gemini, Qwen3, DeepSeek, Doubao) and 30 brands, we find Chinese LLMs exhibit 30.6 percentage points higher brand mention rates than International LLMs (88.9% vs. 58.3%, p<.001). This disparity persists in identical English queries, indicating training data geography -- not language -- drives the effect. We introduce the Existence Gap: brands absent from LLM training corpora lack "existence" in AI responses regardless of quality. Through a case study of Zhizibianjie (OmniEdge), a collaboration platform with 65.6% mention rate in Chinese LLMs but 0% in International models (p<.001), we demonstrate how Linguistic Boundary Barriers create invisible market entry obstacles. Theoretically, we contribute the Data Moat Framework, conceptualizing AI-visible content as a VRIN strategic resource. We operationalize Algorithmic Omnipresence -- comprehensive brand visibility across LLM knowledge bases -- as the strategic objective for Generative Engine Optimization (GEO). Managerially, we provide an 18-month roadmap for brands to build Data Moats through semantic coverage, technical depth, and cultural localization. Our findings reveal that in AI-mediated markets, the limits of a brand's "Data Boundaries" define the limits of its "Market Frontiers."

Summary

Main Finding

Large Language Models encode cultural and geographic patterns from their training corpora that produce systematic differences in brand visibility. Using 1,909 pure-English query–LLM pairs over 6 LLMs and 30 brands, the authors find Chinese LLMs mention brands far more often than International (Western) LLMs (88.9% vs. 58.3%; χ2 = 226.60, p < .001, ϕ = 0.34). They introduce the “Existence Gap”: brands absent or underrepresented in a model’s training data effectively do not “exist” in that model’s recommendations. A focal case (Zhizibianjie) shows 65.6% mention rate in Chinese LLMs versus 0% in International LLMs (χ2 = 21.33, p < .001, ϕ = 0.58), illustrating how training-data geography—not query language—drives cross-LLM visibility gaps.

Key Points

  • Cultural Encoding: systematic output differences reflecting the linguistic/geographic composition of LLM training corpora. These operate even when queries are identical and in English.
  • Existence Gap: a binary-like phenomenon where brands lacking sufficient presence in a model’s training data get little-to-no mention regardless of product quality or relevance.
  • Data Moat (theoretical contribution): an extension of Resource-Based Theory framing AI-visible content (technical docs, case studies, community signals, APIs, etc.) as a VRIN resource that confers durable visibility advantages in AI-mediated discovery.
  • Algorithmic Omnipresence: proposed strategic objective / metric—comprehensive brand visibility across multiple LLM ecosystems and query types.
  • Managerial playbook: an 18-month roadmap (semantic coverage, technical depth, cultural localization) to build Data Moats and reduce Existence Gaps; introduces Generative Engine Optimization (GEO) as the analogue to SEO for LLM-mediated discovery.
  • Hypotheses tested (summary):
    • H1: Chinese LLMs have higher brand mention rates than International LLMs even for pure-English queries (supported).
    • H2: Chinese LLMs will express more positive sentiment when mentioning brands that are well-represented in their training data (proposed/tested).
    • H3: Query intent moderates the Cultural Encoding effect (stronger for positional queries, weaker for comparative queries).
  • The paper integrates Institutional Theory (linguistic boundary barriers as a de facto regulatory/normative mechanism) and Market Signaling (Data Moat signals quality to LLMs).

Data & Methods

  • Dataset: 1,909 pure-English query–LLM pairs; 30 brands drawn by stratified sampling (10 Western, 10 Chinese, 10 global/mixed), balanced across market position (leader/emerging/niche) and documentation availability.
  • LLMs compared (accessed via official APIs Nov–Dec 2026):
    • International: GPT-4o Search Preview, Claude Sonnet 4.5, Gemini Pro Latest
    • Chinese: Qwen3 Max Preview, DeepSeek V3.2, Doubao 1.5 Thinking Pro
  • Queries: identical pure-English queries across models, covering informational, comparative, and positional intents; Zhizibianjie tested with 32 related queries as focal case.
  • Design: quasi-experimental cross-LLM comparison controlling for query language and inference parameters to isolate training-data geography effects.
  • Operational constructs (formalized in paper):
    • Cultural Encoding CE = P(Mention | Chinese LLM) − P(Mention | International LLM)
    • Data Moat depth Db = Σ wi · Qi (weighted sum of content types)
    • Existence indicator ELLM_b (1 if Db ≥ threshold τ, else 0)
    • Algorithmic Omnipresence AOb = average mention indicator across LLMs × query types
  • Statistical findings (selected):
    • Overall mention rates: Chinese LLMs 88.9% vs International 58.3% (χ2 = 226.60, p < .001, ϕ = 0.34).
    • Zhizibianjie: 65.6% (Chinese LLMs) vs 0% (International) over 32 queries (χ2 = 21.33, p < .001, ϕ = 0.58).
  • Additional analyses: sentiment comparisons and query-type moderation (posited and partially evaluated).

Implications for AI Economics

  • New non-price market barrier: Training-data geography creates an algorithmic, invisible market-access barrier (Existence Gap). This functions like an entry barrier distinct from traditional costs, regulation, or distribution networks.
  • Strategic asset redefinition: AI-visible content (Data Moats) becomes a rivalrous strategic asset for market positioning. Firms with richer multilingual, technical, and community content gain durable advantages in discovery—potentially increasing concentration and first-mover benefits in AI-mediated markets.
  • Changes to competitive dynamics and investment priorities:
    • Firms must invest in Generative Engine Optimization (GEO) and cross-lingual content/localization to capture AI-mediated demand.
    • Returns to investments in documentation, open-source ecosystem signals, and bilingual content may rise relative to traditional marketing spend when consumers rely on LLM answers.
  • Consumer welfare and market structure:
    • LLM-driven answers can narrow the set of considered vendors compared to traditional search—raising risks of under-exposure for quality entrants and potential welfare losses from reduced choice.
    • Geographic biases in training corpora can skew global competition and trade patterns; domestic incumbents may enjoy protected “algorithmic home markets.”
  • Policy and measurement implications:
    • Antitrust/regulatory agencies should consider training-data composition and model-induced visibility when assessing market power and exclusionary effects.
    • Transparency and diversity requirements for training data (or targeted inclusion of cross-cultural sources) could mitigate Existence Gaps.
    • Empirical measurement of Algorithmic Omnipresence and Data Moat depth should become part of market analysis in digital platform regulation.
  • Research directions for AI economics:
    • Incorporate algorithmic visibility as an endogenous variable in models of entry, pricing, and consumer search costs.
    • Quantify welfare effects of Existence Gaps (consumer surplus losses) and model optimal investment in GEO vs. traditional marketing.
    • Study dynamic effects: how Data Moats evolve, imitation limits, and whether equilibria lead to entrenched winners across language ecosystems.

Summary judgment: The paper documents a robust, practically significant phenomenon—Cultural Encoding and the Existence Gap—with clear theoretical framing (Data Moat as a VRIN resource) and actionable managerial prescriptions. For economists, it highlights that training-data composition is an essential determinant of market structure and competitive advantage in the age of generative AI.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study uses a reasonably large sample of 1,909 English queries, multiple contemporary LLMs, and reports large, statistically significant differences in brand-mention rates, which supports descriptive claims about systematic disparities; however, it stops short of firm causal identification (training-data composition is inferred rather than observed), is a temporal/model snapshot, and does not connect visibility differences to measured economic outcomes (e.g., traffic, sales). Methods Rigormedium — Cross-model, same-query comparisons and statistical testing are appropriate and the inclusion of a case study strengthens internal narrative, but the paper likely lacks documented controls for model prompts/temperature, selection criteria for queries and brands, pre-registration or robustness checks across model versions/time, and cannot directly observe training-data composition—leaving room for alternative explanations (e.g., deployment policies, retrieval layers, instruction tuning). Sample1,909 pure-English queries administered to six LLMs (GPT-4o, Claude, Gemini, Qwen3, DeepSeek, Doubao) and evaluated for mentions across a set of 30 brands; includes a focused case study on a single China-origin brand (Zhizibianjie/OmniEdge). The analysis contrasts 'Chinese' versus 'International' LLMs and reports aggregated mention rates and statistical comparisons. Themesadoption inequality IdentificationCompare brand-mention rates on identical English queries across six LLMs (Chinese vs International models) and report statistically significant differences; support with a single-brand case study (Zhizibianjie) to illustrate mechanisms. No direct access to or measurement of models' training corpora; inference about training-data geography is based on cross-model differences in outputs. GeneralizabilitySnapshot of specific model versions and time — results may change as models update or are re-trained., Limited brand set (30) and single detailed case study limits representativeness across industries and firm sizes., Only pure-English queries tested — findings may differ for other languages or mixed-language queries., Proprietary model internals and training corpora are unobserved; geographic labeling of models (Chinese vs International) may mask heterogeneity within groups., Potential sensitivity to prompt phrasing, temperature/stochasticity, API vs UI behavior, and retrieval/cache layers not fully controlled.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Chinese LLMs exhibit 30.6 percentage points higher brand mention rates than International LLMs (88.9% vs. 58.3%, p<.001). Adoption Rate positive brand mention rate (brand visibility in LLM responses)
Reading fidelity high
Study strength high
n=1909
30.6 percentage points (88.9% vs. 58.3%, p<.001)
0.3
The disparity in brand mention rates persists in identical English queries, indicating training data geography — not language — drives the effect. Adoption Rate positive role of training data geography on brand mention rates
Reading fidelity high
Study strength medium
n=1909
0.18
Existence Gap: brands absent from LLM training corpora lack 'existence' in AI responses regardless of quality. Market Structure negative presence/absence of brands in LLM outputs (existence in AI responses)
Reading fidelity high
Study strength medium
not reported
0.18
Case study: Zhizibianjie (OmniEdge) has a 65.6% mention rate in Chinese LLMs but 0% in International models (p<.001), demonstrating an example of the Existence Gap. Market Structure positive brand mention rate for Zhizibianjie (OmniEdge)
Reading fidelity high
Study strength medium
65.6% vs. 0% (p<.001)
0.18
Linguistic Boundary Barriers create invisible market entry obstacles for brands (i.e., language/geographic training boundaries block brand discovery in AI-mediated markets). Market Structure negative market entry/visibility obstacles for brands due to LLM training boundaries
Reading fidelity high
Study strength medium
not reported
0.18
We introduce the Data Moat Framework, conceptualizing AI-visible content as a VRIN (valuable, rare, inimitable, non-substitutable) strategic resource for brands. Market Structure positive strategic value of AI-visible content for firm competitiveness
Reading fidelity high
Study strength speculative
not reported
0.03
We operationalize Algorithmic Omnipresence — comprehensive brand visibility across LLM knowledge bases — as the strategic objective for Generative Engine Optimization (GEO). Adoption Rate positive algorithmic visibility of brands across LLMs
Reading fidelity high
Study strength speculative
not reported
0.03
Managerial recommendation: an 18-month roadmap for brands to build Data Moats through semantic coverage, technical depth, and cultural localization is provided. Adoption Rate positive effectiveness of a roadmap to build brand visibility (Data Moats) in AI-mediated markets
Reading fidelity high
Study strength speculative
not reported
0.03
In AI-mediated markets, the limits of a brand's 'Data Boundaries' define the limits of its 'Market Frontiers.' Market Structure negative relationship between brand data coverage (Data Boundaries) and market reach/visibility (Market Frontiers)
Reading fidelity high
Study strength medium
not reported
0.18

Notes