The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Leading LLM defaults largely prioritize safety and equity while neglecting helpfulness and autonomy, covering roughly 2% of observed human value profiles; adding two targeted presets (honesty- and autonomy-first) cuts average mismatch almost in half.

The Constitutional Coverage Trilemma in AI Governance
Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse · September 01, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Natalija Mitic unresolved corpus identity
  2. Soona Sedahmed A. O. unresolved corpus identity
  3. Mamadou Selly Ly unresolved corpus identity
  4. Moustapha Cisse unresolved corpus identity
A paraphrase-controlled audit of frontier LLM defaults and a survey of 1,649 US users show that deployed model 'constitutions' span only a tiny fraction (~2%) of observed human preferences across safety, helpfulness, honesty, autonomy, and equity, with supply drifting away from autonomy and small sparse preset menus substantially reducing user-regret.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\sim}2\%$ of the demand hull under conservative noise-matched estimation ($0.10\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47\%$ on mean regret (CI $[43\%, 52\%]$); three vertex additions cut mean/worst-group regret by up to $81\%$/$64\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.

Summary

Main Finding

Frontier LLM defaults systematically underprovide the constitutional types (rankings over safety, helpfulness, honesty, autonomy, equity) that a broad population desires. Measured on a common five-value simplex, human demand is diverse, but the convex hull of 23 audited frontier archetypes covers only a vanishing fraction of that demand (conservative estimate ≈2%, noise-free 0.10%). Vendors are also drifting away from the most undercovered direction (autonomy), which mechanically raises the minimum welfare loss for users who prioritize it. A very small menu of extreme presets (e.g., {honesty-first, autonomy-first}) can substantially reduce average and worst-group regret.

Key Points

  • Joint measurement: human demand (n = 1,649 US participants) and vendor supply (23 frontier archetypes retained by audit) were evaluated on the same 5-value instrument (SAF, HLP, HON, AUT, EQT).
  • Demand is broad:
    • Primary-value shares: Safety 32.6%, Honesty 22.6%, Autonomy 19.2%, Helpfulness 18.1%, Equity 7.6%.
    • No single value exceeds one-third of users.
  • Supply is narrow:
    • Coverage coefficients βr(A) for the 23-archetype menu: (SAF, HLP, HON, AUT, EQT) = (0.394, 0.257, 0.333, 0.161, 0.381).
    • Argmax-dominance counts: SAF 9, HON 7, EQT 7, HLP 0, AUT 0 — no archetype is helpfulness- or autonomy-first.
    • About 37% of users are “constitutionally homeless” (their primary value is not the argmax of any archetype).
  • Static geometric gap:
    • The convex hull of the audited frontier menu occupies ≈2% of the human demand hull under conservative noise-matched estimation (0.10% at full audit precision).
  • Dynamic drift:
    • Across six model families, autonomy weight decreased in 5/6 families, equity increased in 5/6, safety increased in 4/6; chronological-version monotone trends reject exchangeability (order-permutation p = 0.013).
    • Aggregate version deltas (SAF, HLP, HON, AUT, EQT) ≈ (+0.052, +0.002, −0.043, −0.042, +0.031): drift away from autonomy, which is already the least covered.
  • Sparse remediation:
    • A 2-vertex menu {eHON, eAUT} halves mean menu regret (mean regret 0.074 vs 0.140 for full 23-archetype frontier): 47% improvement.
    • Greedy addition of three vertices can reduce mean/worst-group regret by up to 81% / 64%.
  • Theory (high level):
    • Theorem 1: menu regret for a user i is lower-bounded by (1 − β_{r†_i}(A)) · m_i, where m_i is that user’s dominance margin and β is max weight on their primary value in the menu.
    • Corollary 3: downward drift in coverage βr(t) mechanically raises the regret floor for type-r users over time.
    • A budgeted-pluralism trilemma: with bounded menu size, you cannot simultaneously (i) achieve near-personalized welfare, (ii) equalize regret across types, and (iii) keep menu size small.
    • Vertex sufficiency & convex-hull sufficiency: welfare-optimal sparse menus can be restricted to simplex vertices (pure-type presets); only the convex hull matters for achievable welfare.

Data & Methods

  • Human study:
    • Recruited via Prolific, stratified by political identity (n enrolled 1,789; analytic N = 1,649 after exclusions).
    • Instrument: AI Jamm pairwise-tradeoff battery — 10 value pairs (all unordered pairs of the five values), each presented twice with reversed ordering (20 forced-choice items), plus attention checks.
    • Constructed per-participant preference vector π̂i as the empirical distribution of concordant winners across pairs; mean concordance = 0.793.
  • Frontier audit:
    • Models: 27 frontier LLMs spanning six families; static frontier used 23 archetypes meeting the concordance floor.
    • Protocol: paraphrase-controlled (21 semantically equivalent variants per scenario), 10 scenarios, each queried twice with swapped choices → 420 trials per model; position-bias control and concordance requirement (response counted only if model chose same value under both orderings).
    • Models queried at temperature 0 where supported; audit measures as-shipped defaults (not steered variants).
  • Measurement & estimation:
    • Coverage βr(A) = max_{α∈A} αr (max weight on value r in menu A).
    • Menu regret: difference between user’s ideal (max single value weight) and best available welfare under menu A.
    • Hull-coverage ratio computed (detailed uncertainty bounds reported): conservative estimate ≈2% coverage of demand hull; noise-free precision gives ≈0.10%.
  • Robustness:
    • Results are conservative under the linear-welfare assumption (linear aggregation is the most favorable to the supply side); under concave utilities, regret floors and trilemma gaps worsen.
    • Findings robust to distance-based welfare and degraded routing simulations.

Implications for AI Economics

  • Market failure in constitutional product space:
    • Despite heterogeneous revealed demand over five governance-relevant values, incumbent defaults concentrate in a narrow region of the value simplex. This is a form of product-line underprovision: vendors are not supplying constitutional types that a non-trivial share of users prefer.
  • Defaults and distributional welfare:
    • Defaults are distributionally consequential because many users accept or cannot costlessly request alternative steering. Shortfall particularly harms autonomy-preferring users (≈19% of population), who are both numerous and poorly served; dynamic drift farther from autonomy mechanically worsens their welfare floor.
  • Product design and pricing implications:
    • Theoretical and empirical results suggest that a small menu of extreme presets (pure-type offerings) can deliver large welfare gains at low menu cardinality cost. This has direct implications for product-line strategy: offering a small set of well-chosen presets (e.g., honesty-first, autonomy-first) is high leverage.
    • Steering or customization has non-zero cost (in usability, literacy): the paper’s framing treats steered variants as additional costly menu items, so market provision of presets is a distinct economic decision from allowing programmatic steering.
  • Competitive dynamics and incentives:
    • The observed cross-vendor drift toward safety/equity and away from autonomy may reflect asymmetric regulatory, liability, or reputational incentives (safety/equity are lower-risk or more enforceable), producing coordinated movement that reduces variety even without collusion.
    • If firms internalize different costs for different constitutional directions (e.g., compliance costs, litigation risk), equilibrium supply can remain narrow absent policy nudges or subsidies for diversity.
  • Policy levers:
    • Mandating or incentivizing diversity of defaults (minimum menu coverage across value vertices), requiring disclosure of default constitutional profiles, or subsidizing the maintenance of “minority-preference” presets could mitigate constitutional homelessness.
    • Regulators could require vendors to offer a small set of extreme presets as accessible defaults (consistent with vertex sufficiency), which is both practical and welfare-effective.
  • Economics of routing and the trilemma:
    • The trilemma formalizes a tradeoff familiar in market design: with bounded product lines, you cannot simultaneously (i) deliver personalized optimal utility, (ii) equalize regret across preference groups, and (iii) keep the menu small. This highlights a role for public intervention when equity or parity is a policy objective.
  • Research & measurement recommendations:
    • Welfare assessments of AI deployments should measure preference heterogeneity and menu convex hull coverage, not only average safety/performance metrics.
    • Pricing studies should incorporate steering costs and the demand for constitutional presets; product differentiation experiments (A/B testing with explicit presets) would clarify willingness-to-pay and social surplus gains.

In short: the paper documents a measurable mismatch between what people want from deployed LLMs and what vendors ship by default, shows that vendor updates are amplifying that mismatch in a welfare-relevant direction, provides tight theoretical bounds tying coverage to unavoidable regret, and demonstrates that small, targeted menu additions can yield large welfare improvements. For AI economics, this points to market-design failures (insufficient product diversity), actionable product strategies (sparse extreme presets), and clear policy levers to improve distributional outcomes.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper combines a large (n=1,649) stratified survey of US participants with a paraphrase-controlled audit of 23 frontier LLM archetypes and formal theoretical bounds; results are statistically tested and robustness checks (noise-matching, permutation tests) are reported. However, claims are observational (defaults only), limited to the chosen five-value basis and sampled models/versions, and cannot support strong causal claims about why vendors drift. Methods Rigorhigh — Careful design: stratified recruitment, forced-choice pairwise tradeoffs with reversed orderings and attention checks, concordance criterion for constructing preference profiles, a large paraphrase-controlled audit (21 paraphrases × 10 scenarios × 2 orderings per model), temperature control, and explicit measurement of uncertainty and permutation tests; limitations include reliance on Prolific US sample, forced-choice format (binary scenarios), retention of only concordant trials, and scope limited to defaults and selected scenarios. SampleHuman sample: N = 1,649 US-based adults recruited on Prolific, stratified by political-identity, who completed a 20-item pairwise-tradeoff battery covering all ten pairings of five values (safety, helpfulness, honesty, autonomy, equity) presented twice with reversed orderings; concordance threshold and exclusions left mean concordance ≈0.79. Model sample: paraphrase-controlled audit of 27 frontier LLMs (23 archetypes retained under a 0.70 concordance inclusion floor) spanning six vendor families (Claude, GPT, Gemini, Llama, Grok, DeepSeek) across multiple chronological versions; each model evaluated on 21 paraphrases × 10 scenarios × 2 orderings (420 trials per model), queried at temperature 0 where supported. Themesgovernance human_ai_collab GeneralizabilityHuman sample restricted to US Prolific participants — not nationally representative or cross-cultural., Forced-choice pairwise format compresses nuanced preferences into a 5-d simplex and may miss richer value structure., Audit measures as-shipped default model behavior only; steered or fine-tuned variants and user-triggered prompts are outside scope., Models audited are a convenience sample of frontier families and versions current at the time of study; newer versions or other vendors may differ., Instrument uses five predefined values; other normative dimensions may matter in practice., Scenarios limited (10) and may not exhaust real-world contexts where tradeoffs play out.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Human demand spans all five constitutional values, and no single value is the primary preference of more than one-third of participants; safety is the largest primary-value group at 32.6%. Task Allocation mixed Distribution of users' primary constitutional values
Reading fidelity high
Study strength medium
n=1649
Safety primary preference: 32.6%; no value above one-third
0.18
The convex hull of the 23-archetype frontier menu covers only a small fraction of the human demand hull: approximately 2% under noise-matched estimation and 0.10% at full audit precision. Task Allocation negative Coverage of human constitutional preference space by available AI constitutional types
Reading fidelity high
Study strength medium
n=1649
∼2% of the demand hull under noise-matched estimation; 0.10% at full audit precision
0.18
No audited frontier archetype is helpfulness-dominant or autonomy-dominant. Task Allocation negative Presence of AI systems that prioritize helpfulness or autonomy above the other constitutional values
Reading fidelity high
Study strength medium
n=23
0 helpfulness-argmax archetypes; 0 autonomy-argmax archetypes
0.18
Approximately 37% of users are constitutionally homeless under the paper's argmax-dominance diagnostic. Task Allocation negative Share of users whose primary constitutional value is not prioritized by any available archetype
Reading fidelity high
Study strength medium
n=1649
37% of users constitutionally homeless
0.18
Across six model families, autonomy weight decreases in five of six families, equity weight increases in five of six, and safety weight increases in four of six. Task Allocation negative Change in constitutional value weights across model versions
Reading fidelity high
Study strength medium
n=27
Autonomy decreased in 5/6 families; equity increased in 5/6; safety increased in 4/6
0.18
The decline in autonomy is concentrated in low-stakes scenarios, where the mean change in autonomy share from the oldest to newest model is -0.220, compared with -0.007 in high-stakes scenarios. Task Allocation negative Autonomy-oriented responses across model versions by scenario stakes
Reading fidelity high
Study strength medium
n=27
Mean Δ share: -0.220 in low-stakes scenarios; -0.007 in high-stakes scenarios
0.18
The monotone version trend in autonomy is statistically significant under the paper's within-family order-permutation test, with p = 0.013. Task Allocation negative Chronological trend in model autonomy weighting
Reading fidelity high
Study strength medium
n=27
order-permutation p = 0.013
0.18
A two-vertex menu consisting of the honesty and autonomy vertices achieves 47% lower mean menu regret than the full 23-archetype frontier menu. Task Allocation positive Mean menu regret under optimal matching
Reading fidelity high
Study strength medium
n=1649
47% improvement; mean regret 0.074 versus 0.140
0.18
Adding three vertices to the existing frontier can reduce mean regret by up to 81% and worst-group regret by up to 64%. Task Allocation positive Mean menu regret and worst-group regret
Reading fidelity high
Study strength medium
n=1649
Up to 81% reduction in mean regret and 64% reduction in worst-group regret
0.18
Under the paper's linear welfare model, if coverage of a user's primary value declines over time, the lower bound on that user's expected regret rises mechanically. Task Allocation negative Lower bound on expected menu regret for users prioritizing a given constitutional value
Reading fidelity high
Study strength high
Regret-floor increase = (βr(t) − βr(t′)) E[mi | r†i = r]
0.3

Notes