The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Shared language models push authors toward a single writing norm, but per-user personalization can preserve stylistic variety; moreover, individually rational conformity to LLM suggestions can over-standardize language and impose a social cost that may become large in some regimes.

Linguistic Monoculture in LLM-Assisted Language Use
Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen · July 29, 2026
arxiv theoretical medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Suhas Thejaswi unresolved corpus identity
  2. Juhi Kulshreshta unresolved corpus identity
  3. Lutz Oettershagen unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Suhas Thejaswi provider ID
  2. Juhi Kulshreshta provider ID
  3. Lutz Oettershagen provider ID
A formal model and simulations show that shared LLM assistance can drive population-level linguistic convergence (monoculture), recursive retraining shifts the shared norm without necessarily changing pairwise spread, and personalization can preserve stylistic diversity, while individually optimal conformity can exceed the social optimum, creating a quantifiable 'price of monoculture.'

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.

Summary

Main Finding

The paper formalizes how repeated interactions between human authors and LLMs can reduce population-level linguistic variation ("linguistic monoculture"), characterizes when and how that homogenization happens, and shows personalization and heterogeneous preferences can preserve diversity. It proves exponential convergence under a range of deployment regimes, identifies the roles of conformity, learning/adaptation and recursive retraining, and shows individually optimal conformity can exceed the social optimum (a negative externality captured as a "price of monoculture") — finite in many instances but unbounded in certain parameter regimes.

Key Points

  • Model objects: authors and LLM outputs are represented as probability distributions over a common set of linguistic features F. Diversity is measured by average pairwise Jensen–Shannon (JS) divergence Dt; alignment author–model by Mt; diversity among personalized models by Qt.
  • Three interaction mechanisms (IMs):
  • IM1 — Shared model with fixed distribution q0: all authors adapt toward the same fixed LLM output.
  • IM2 — Shared model with recursive updates: a single model is retrained on author outputs (population-average feedback).
  • IM3 — Personalized models with recursive updates: each author has a personalized model updated from their own outputs plus population feedback; ρ parameterizes degree of personalization.
  • Individual update form (reduced form): pt+1_i = (1 − αi) pt_i + αi Ai(qt_i, zt_i), where αi is adaptation rate and Ai mixes author preferences and model suggestion (e.g., conformity weight λi). Model retraining (shared): qt+1 = β qt + (1 − β) average_author_outputs. Personalized model update mixes author-specific and population feedback (parameters β, ρ).
  • Main analytic results:
    • Convergence: all three IMs converge to equilibria; convergence is exponential with rates depending on α, β and conformity parameters.
    • IM1 (fixed shared model): with homogeneous adaptation, diversity decays exponentially at a rate set by min αi; with author-specific fixed adaptation targets ai, each author converges to ai and diversity converges to the diversity among {ai}.
    • Conformity parameter λ interpolates between full adoption of q0 (λ=1 → monoculture) and preservation of preferred styles (λ=0 → persistence of diversity). Limiting diversity scales as O(1−λ).
    • IM2 (shared recursive): recursion relocates the equilibrium model distribution (q*) to an endogenous value determined by authors' preferences and weights, but under common conformity the pairwise geometry of the author cloud is unchanged relative to IM1 — so recursion anchors the cloud without necessarily changing pairwise separations.
    • IM3 (personalization): personalized recursive updates can produce distinct author–model equilibria; equilibrium author and model distributions retain author-specific components (proportional to personalization parameter ρ and other factors). Thus personalization can sustain nonzero Dt and Qt.
    • Comparative geometry: under common λ, IM1 and IM2 have identical limiting pairwise differences; IM3 inflates pairwise distances by factor 1/(1 − ρλ) (so ρλ > 0 increases pairwise separation).
  • Strategic conformity and externality:
    • Authors choose conformity (λ) to trade off private benefits (clarity/legibility/reward) versus distinctiveness preferences.
    • Individually optimal conformity generally exceeds the social optimum because authors do not internalize the value their distinctiveness provides others — this creates a negative externality (price of monoculture).
    • The authors derive bounds on this price in symmetric regimes and identify parameter regimes where the price can diverge (notably when distinctiveness valuation dominates authenticity incentives).
  • Simulations: synthetic experiments illustrate how IM1/IM2/IM3 produce differing long-run diversity patterns and validate analytic comparisons.

Data & Methods

  • Nature of study: theoretical / formal-analytic model + synthetic simulations. No empirical dataset of real texts is used.
  • Mathematical representation:
    • Linguistic style = distribution over m features (vector in simplex Δm−1).
    • Diversity = average pairwise Jensen–Shannon divergence Dt = (1/(n(n−1))) Σ_{i≠j} JS(p_i, p_j).
    • Quadratic, translation-invariant proxy D^ used to isolate pairwise geometry: average squared L2 distances.
  • Dynamics specified by linear mixture updates for authors and models with parameters:
    • αi ∈ (0,1): author adaptation speed
    • λi ∈ [0,1]: conformity weight toward model vs. personal preference ri
    • β ∈ [0,1]: model retention when retraining
    • ρ ∈ [0,1]: degree of personalization in model updates
    • wi: population weights in model retraining
  • Analytical methods:
    • Fixed-point algebra to solve equilibria.
    • Convergence proofs giving exponential bounds (e.g., for homogeneous adaptation in IM1, to get Dt ≤ ε need t ≥ log(2/ε)/α_min).
    • Comparative statics with respect to α, λ, β, ρ, and weight structure.
    • Price-of-monoculture analysis: utility model where conformity yields private benefits; social welfare includes value of distinctiveness to others; compute individually optimal vs socially optimal λ and bound loss.
  • Simulations: controlled synthetic initializations and parameter sweeps to illustrate equilibrium diversity under different IMs.
  • Key assumptions/limitations:
    • Reduced-form abstraction: models and authors are distributions over features (no semantics or deeper cognitive modeling).
    • Simplified linear-mixture update rules; adaptation operators often assumed time-invariant or of simple convex-mixture form.
    • Jensen–Shannon chosen as divergence metric; results on pairwise geometry use L2-quadratic proxy to separate translation effects.
    • Externalities characterized in stylized symmetric regimes; real-world calibration not provided.

Implications for AI Economics

  • Externalities from shared LLMs: widespread use of a common LLM (especially if retrained on generated or assisted outputs) creates a negative externality by reducing the "product variety" of linguistic styles that others value. This is economically analogous to algorithmic monoculture in decision-making domains.
  • Market and platform design levers:
    • Personalization (higher ρ) can mitigate monoculture by preserving author-specific equilibria; platforms could offer or default to personalized models to maintain diversity.
    • Retaining separation between human-produced text and model training pipelines (reducing recursive feedback, increasing β or preventing training on generated text) can prevent endogenous anchoring of norms.
    • Product differentiation (multiple competing models, different decoding/temperature regimes, style-aware prompts) increases choice and may reduce the social loss from homogenization.
  • Incentive and policy responses:
    • Institutions that reward legibility (journals, employers, reviewers) can unintentionally push authors toward conformity; adjusting evaluation incentives to value stylistic distinctiveness or penalize over-reliance on a single model could reduce the externality.
    • Disclosure policies or provenance labels could help recipients adjust valuation of distinctiveness and potentially internalize some social value.
    • Subsidies or reputation mechanisms for distinctive voices (or penalties for excessive conformity) could align private incentives with social welfare.
  • Strategic considerations for platform/providers:
    • Providers optimizing only for aggregate user satisfaction (legibility/fluency) may accelerate monoculture; internalizing social value of diversity suggests offering features (personalization, control knobs, mixed-sampling) or limiting automatic retraining on assisted outputs.
    • The paper shows cases where the social loss (price of monoculture) can be large or unbounded in extreme parameter regimes — sectors where distinctiveness is particularly valuable (creative writing, niche scientific communication, cultural expression) may need stronger protective measures.
  • Empirical agenda for AI economics:
    • Quantify real-world λ (conformity incentives) and how institutional rewards shape it across domains (science, journalism, education).
    • Measure how retraining on generated/assisted text shifts population-level feature distributions over time.
    • Assess welfare trade-offs between clarity/legibility benefits from standardization and diversification externalities on discovery, signaling, and cultural value.
  • Takeaway for policy: interventions (personalization, training-data governance, incentive design) can materially affect population-level linguistic diversity; their costs and benefits should be evaluated in contexts where stylistic distinctiveness carries social value.

If you want, I can (a) translate the key propositions into compact formulas for a slide or policy brief, (b) produce example parameter regimes illustrating when personalization is most effective, or (c) sketch empirical designs to estimate the conformity parameter λ and measure real-world price-of-monoculture.

Assessment

Paper Typetheoretical Evidence Strengthmedium — The paper presents rigorous mathematical propositions, proofs, and synthetic simulations that demonstrate possible dynamics; however it contains no real-world empirical validation or causal estimates using observational or experimental data, limiting claims about real systems. Methods Rigorhigh — The authors develop a clear reduced-form, formal model with provable convergence results, bounds on rates, and distinct interaction mechanisms; assumptions and limiting cases are stated and multiple mechanisms (fixed/shared, recursive, personalized) are analyzed, though the reduced-form abstractions trade realism for analytical tractability. SampleNo empirical sample; analytic model of n authors whose linguistic styles are distributions over m features (probability simplex), coupled to either a shared fixed model, a shared recursively-updated model, or personalized recursively-updated models; parameters include adaptation rates (α_i), conformity weights (λ_i), model retention (β), and personalization weight (ρ); results illustrated via synthetic simulations (parameter sweeps/paired runs). Themeshuman_ai_collab governance GeneralizabilityNo empirical validation — conclusions are model-driven and may not map quantitatively to real author–LLM systems., Reduced-form representation abstracts away semantic/content changes and models style only as distributions over pre-specified features., Assumes particular update rules (linear mixtures, weighted averages) and Jensen–Shannon diversity metric; other behavioral update forms may yield different dynamics., Ignores institutional heterogeneity, strategic platform incentives, differential access to personalization, and real-world training pipelines' complexities., Synthetic simulations depend on parameter choices and initializations that may not reflect deployed LLM behavior or editorial incentives.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Under a fixed shared model and homogeneous author adaptation, population-level linguistic diversity decays exponentially toward zero. Other negative Average pairwise Jensen–Shannon divergence among author linguistic-style distributions
Reading fidelity high
Study strength high
t ≥ log(2/ε) / αmin guarantees D_t ≤ ε
0.2
With a fixed shared model but author-specific, time-invariant adaptation targets, each author converges to their own target, and long-run population diversity equals the diversity among those targets. Other positive Long-run population-level linguistic diversity
Reading fidelity high
Study strength high
not reported
0.2
Under a fixed shared model, conformity to a common model-induced norm reduces long-run linguistic diversity; with a common conformity level λ, limiting diversity is of order O(1 − λ), and full conformity causes collapse to the shared model distribution. Other negative Limiting average pairwise linguistic diversity
Reading fidelity high
Study strength high
D∞(λ) = O(1 − λ)
0.2
Recursive feedback alone does not force linguistic monoculture when authors are pulled toward fixed, author-specific adaptation targets. Other null_result Population-level linguistic diversity under recursive model updates
Reading fidelity high
Study strength high
not reported
0.2
For a recursively updated shared model with common conformity, recursion changes the equilibrium’s location by moving authors toward an endogenous shared norm, but it does not change pairwise author separation in the quadratic diversity measure. Other null_result Pairwise geometric diversity among author linguistic distributions
Reading fidelity high
Study strength high
D̂∞(IM 2) = D̂∞(IM 1)
0.2
Personalization can preserve nonzero diversity by producing a family of distinct author–model equilibria when the personalization parameter ρ is greater than zero; when ρ = 0, all personalized models collapse to the same population-level distribution. Other mixed Author linguistic diversity and diversity among personalized model distributions
Reading fidelity high
Study strength high
not reported
0.2
Under common conformity, personalization increases pairwise quadratic diversity relative to shared-model mechanisms by the factor 1/(1 − ρλ), with equality only when ρλ = 0. Other positive Limiting translation-invariant quadratic diversity among author distributions
Reading fidelity high
Study strength high
D̂∞(IM 3) = D̂∞(IM 1) / (1 − ρλ)^2
0.2
Individually rational authors may choose more conformity than is socially optimal because they do not internalize the value that their distinctiveness provides to others, creating a negative externality and a finite but potentially unbounded price of monoculture. Other negative Social value of linguistic distinctiveness and loss of linguistic diversity
Reading fidelity high
Study strength medium
not reported
0.12

Notes