Leading LLM defaults largely prioritize safety and equity while neglecting helpfulness and autonomy, covering roughly 2% of observed human value profiles; adding two targeted presets (honesty- and autonomy-first) cuts average mismatch almost in half.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\sim}2\%$ of the demand hull under conservative noise-matched estimation ($0.10\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47\%$ on mean regret (CI $[43\%, 52\%]$); three vertex additions cut mean/worst-group regret by up to $81\%$/$64\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.
Summary
Main Finding
Frontier LLM defaults systematically underprovide the constitutional types (rankings over safety, helpfulness, honesty, autonomy, equity) that a broad population desires. Measured on a common five-value simplex, human demand is diverse, but the convex hull of 23 audited frontier archetypes covers only a vanishing fraction of that demand (conservative estimate ≈2%, noise-free 0.10%). Vendors are also drifting away from the most undercovered direction (autonomy), which mechanically raises the minimum welfare loss for users who prioritize it. A very small menu of extreme presets (e.g., {honesty-first, autonomy-first}) can substantially reduce average and worst-group regret.
Key Points
- Joint measurement: human demand (n = 1,649 US participants) and vendor supply (23 frontier archetypes retained by audit) were evaluated on the same 5-value instrument (SAF, HLP, HON, AUT, EQT).
- Demand is broad:
- Primary-value shares: Safety 32.6%, Honesty 22.6%, Autonomy 19.2%, Helpfulness 18.1%, Equity 7.6%.
- No single value exceeds one-third of users.
- Supply is narrow:
- Coverage coefficients βr(A) for the 23-archetype menu: (SAF, HLP, HON, AUT, EQT) = (0.394, 0.257, 0.333, 0.161, 0.381).
- Argmax-dominance counts: SAF 9, HON 7, EQT 7, HLP 0, AUT 0 — no archetype is helpfulness- or autonomy-first.
- About 37% of users are “constitutionally homeless” (their primary value is not the argmax of any archetype).
- Static geometric gap:
- The convex hull of the audited frontier menu occupies ≈2% of the human demand hull under conservative noise-matched estimation (0.10% at full audit precision).
- Dynamic drift:
- Across six model families, autonomy weight decreased in 5/6 families, equity increased in 5/6, safety increased in 4/6; chronological-version monotone trends reject exchangeability (order-permutation p = 0.013).
- Aggregate version deltas (SAF, HLP, HON, AUT, EQT) ≈ (+0.052, +0.002, −0.043, −0.042, +0.031): drift away from autonomy, which is already the least covered.
- Sparse remediation:
- A 2-vertex menu {eHON, eAUT} halves mean menu regret (mean regret 0.074 vs 0.140 for full 23-archetype frontier): 47% improvement.
- Greedy addition of three vertices can reduce mean/worst-group regret by up to 81% / 64%.
- Theory (high level):
- Theorem 1: menu regret for a user i is lower-bounded by (1 − β_{r†_i}(A)) · m_i, where m_i is that user’s dominance margin and β is max weight on their primary value in the menu.
- Corollary 3: downward drift in coverage βr(t) mechanically raises the regret floor for type-r users over time.
- A budgeted-pluralism trilemma: with bounded menu size, you cannot simultaneously (i) achieve near-personalized welfare, (ii) equalize regret across types, and (iii) keep menu size small.
- Vertex sufficiency & convex-hull sufficiency: welfare-optimal sparse menus can be restricted to simplex vertices (pure-type presets); only the convex hull matters for achievable welfare.
Data & Methods
- Human study:
- Recruited via Prolific, stratified by political identity (n enrolled 1,789; analytic N = 1,649 after exclusions).
- Instrument: AI Jamm pairwise-tradeoff battery — 10 value pairs (all unordered pairs of the five values), each presented twice with reversed ordering (20 forced-choice items), plus attention checks.
- Constructed per-participant preference vector π̂i as the empirical distribution of concordant winners across pairs; mean concordance = 0.793.
- Frontier audit:
- Models: 27 frontier LLMs spanning six families; static frontier used 23 archetypes meeting the concordance floor.
- Protocol: paraphrase-controlled (21 semantically equivalent variants per scenario), 10 scenarios, each queried twice with swapped choices → 420 trials per model; position-bias control and concordance requirement (response counted only if model chose same value under both orderings).
- Models queried at temperature 0 where supported; audit measures as-shipped defaults (not steered variants).
- Measurement & estimation:
- Coverage βr(A) = max_{α∈A} αr (max weight on value r in menu A).
- Menu regret: difference between user’s ideal (max single value weight) and best available welfare under menu A.
- Hull-coverage ratio computed (detailed uncertainty bounds reported): conservative estimate ≈2% coverage of demand hull; noise-free precision gives ≈0.10%.
- Robustness:
- Results are conservative under the linear-welfare assumption (linear aggregation is the most favorable to the supply side); under concave utilities, regret floors and trilemma gaps worsen.
- Findings robust to distance-based welfare and degraded routing simulations.
Implications for AI Economics
- Market failure in constitutional product space:
- Despite heterogeneous revealed demand over five governance-relevant values, incumbent defaults concentrate in a narrow region of the value simplex. This is a form of product-line underprovision: vendors are not supplying constitutional types that a non-trivial share of users prefer.
- Defaults and distributional welfare:
- Defaults are distributionally consequential because many users accept or cannot costlessly request alternative steering. Shortfall particularly harms autonomy-preferring users (≈19% of population), who are both numerous and poorly served; dynamic drift farther from autonomy mechanically worsens their welfare floor.
- Product design and pricing implications:
- Theoretical and empirical results suggest that a small menu of extreme presets (pure-type offerings) can deliver large welfare gains at low menu cardinality cost. This has direct implications for product-line strategy: offering a small set of well-chosen presets (e.g., honesty-first, autonomy-first) is high leverage.
- Steering or customization has non-zero cost (in usability, literacy): the paper’s framing treats steered variants as additional costly menu items, so market provision of presets is a distinct economic decision from allowing programmatic steering.
- Competitive dynamics and incentives:
- The observed cross-vendor drift toward safety/equity and away from autonomy may reflect asymmetric regulatory, liability, or reputational incentives (safety/equity are lower-risk or more enforceable), producing coordinated movement that reduces variety even without collusion.
- If firms internalize different costs for different constitutional directions (e.g., compliance costs, litigation risk), equilibrium supply can remain narrow absent policy nudges or subsidies for diversity.
- Policy levers:
- Mandating or incentivizing diversity of defaults (minimum menu coverage across value vertices), requiring disclosure of default constitutional profiles, or subsidizing the maintenance of “minority-preference” presets could mitigate constitutional homelessness.
- Regulators could require vendors to offer a small set of extreme presets as accessible defaults (consistent with vertex sufficiency), which is both practical and welfare-effective.
- Economics of routing and the trilemma:
- The trilemma formalizes a tradeoff familiar in market design: with bounded product lines, you cannot simultaneously (i) deliver personalized optimal utility, (ii) equalize regret across preference groups, and (iii) keep the menu small. This highlights a role for public intervention when equity or parity is a policy objective.
- Research & measurement recommendations:
- Welfare assessments of AI deployments should measure preference heterogeneity and menu convex hull coverage, not only average safety/performance metrics.
- Pricing studies should incorporate steering costs and the demand for constitutional presets; product differentiation experiments (A/B testing with explicit presets) would clarify willingness-to-pay and social surplus gains.
In short: the paper documents a measurable mismatch between what people want from deployed LLMs and what vendors ship by default, shows that vendor updates are amplifying that mismatch in a welfare-relevant direction, provides tight theoretical bounds tying coverage to unavoidable regret, and demonstrates that small, targeted menu additions can yield large welfare improvements. For AI economics, this points to market-design failures (insufficient product diversity), actionable product strategies (sparse extreme presets), and clear policy levers to improve distributional outcomes.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Human demand spans all five constitutional values, and no single value is the primary preference of more than one-third of participants; safety is the largest primary-value group at 32.6%. Task Allocation | mixed | Distribution of users' primary constitutional values |
Reading fidelity
high
Study strength
medium
|
n=1649
Safety primary preference: 32.6%; no value above one-third
|
| The convex hull of the 23-archetype frontier menu covers only a small fraction of the human demand hull: approximately 2% under noise-matched estimation and 0.10% at full audit precision. Task Allocation | negative | Coverage of human constitutional preference space by available AI constitutional types |
Reading fidelity
high
Study strength
medium
|
n=1649
∼2% of the demand hull under noise-matched estimation; 0.10% at full audit precision
|
| No audited frontier archetype is helpfulness-dominant or autonomy-dominant. Task Allocation | negative | Presence of AI systems that prioritize helpfulness or autonomy above the other constitutional values |
Reading fidelity
high
Study strength
medium
|
n=23
0 helpfulness-argmax archetypes; 0 autonomy-argmax archetypes
|
| Approximately 37% of users are constitutionally homeless under the paper's argmax-dominance diagnostic. Task Allocation | negative | Share of users whose primary constitutional value is not prioritized by any available archetype |
Reading fidelity
high
Study strength
medium
|
n=1649
37% of users constitutionally homeless
|
| Across six model families, autonomy weight decreases in five of six families, equity weight increases in five of six, and safety weight increases in four of six. Task Allocation | negative | Change in constitutional value weights across model versions |
Reading fidelity
high
Study strength
medium
|
n=27
Autonomy decreased in 5/6 families; equity increased in 5/6; safety increased in 4/6
|
| The decline in autonomy is concentrated in low-stakes scenarios, where the mean change in autonomy share from the oldest to newest model is -0.220, compared with -0.007 in high-stakes scenarios. Task Allocation | negative | Autonomy-oriented responses across model versions by scenario stakes |
Reading fidelity
high
Study strength
medium
|
n=27
Mean Δ share: -0.220 in low-stakes scenarios; -0.007 in high-stakes scenarios
|
| The monotone version trend in autonomy is statistically significant under the paper's within-family order-permutation test, with p = 0.013. Task Allocation | negative | Chronological trend in model autonomy weighting |
Reading fidelity
high
Study strength
medium
|
n=27
order-permutation p = 0.013
|
| A two-vertex menu consisting of the honesty and autonomy vertices achieves 47% lower mean menu regret than the full 23-archetype frontier menu. Task Allocation | positive | Mean menu regret under optimal matching |
Reading fidelity
high
Study strength
medium
|
n=1649
47% improvement; mean regret 0.074 versus 0.140
|
| Adding three vertices to the existing frontier can reduce mean regret by up to 81% and worst-group regret by up to 64%. Task Allocation | positive | Mean menu regret and worst-group regret |
Reading fidelity
high
Study strength
medium
|
n=1649
Up to 81% reduction in mean regret and 64% reduction in worst-group regret
|
| Under the paper's linear welfare model, if coverage of a user's primary value declines over time, the lower bound on that user's expected regret rises mechanically. Task Allocation | negative | Lower bound on expected menu regret for users prioritizing a given constitutional value |
Reading fidelity
high
Study strength
high
|
Regret-floor increase = (βr(t) − βr(t′)) E[mi | r†i = r]
|