The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models cannot reliably simulate intersectional survey respondents: two-feature personas usually reduce to one dominant identity and often ignore race and religion, meaning synthetic samples misrepresent intersectional opinion structure.

Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
Virgile Rennard, Christos Xypolopoulos · August 24, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Virgile Rennard unresolved corpus identity
  2. Christos Xypolopoulos unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Virgile Rennard provider ID
  2. Christos Xypolopoulos unresolved corpus identity
When prompted to simulate intersectional survey respondents, large language models typically collapse multi-attribute personas onto one dominant feature (or at most two), systematically discarding other attributes—notably race and religion—so synthetic samples do not reproduce real intersectional opinion distributions.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.

Summary

Main Finding

Large language models do not reliably simulate intersectional demographic identities. When prompted to answer as members of two- or three-feature demographic subgroups, LLMs typically “collapse” the persona onto one (or at most two) identity dimensions rather than integrating the multiple identities additively as real respondents do. This failure is large, robust across eight models and multiple prompting/readout paradigms, and systematically downweights race and religion—the strongest real drivers of opinion.

Key Points

  • Scope and scale
    • Evaluation against every real intersectional subgroup (n ≥ 20) in 15 waves of Pew’s American Trends Panel (ATP).
    • 15.7 million simulated response distributions in the primary pipeline; 21.1 million across all conditions; eight LLMs tested (7B open models through current frontier models).
  • Core quantitative results
    • For two-feature personas, a single feature’s simulated bias explains the model’s pair output better than the additive combination in 75–82% of subgroups (raw win rates across models).
    • Calibrated “collapse index” (rescaling wins between additive-truth and collapse-truth endpoints) places models about 0.83–0.95 of the way to pure collapse (e.g., GPT-4o-mini ≈ 0.86–0.87).
    • Humans: pair identities are approximately additive. Real subgroups' weights center near full weight on both features (median (α, β) ≈ (0.95, 0.95), total α+β ≈ 1.83).
    • Models: pair biases concentrate on a one-identity budget (median α+β ≈ 0.98), typically dominated by one feature (median dominant share ≈ 0.88).
    • Real intersections grow more distinctive with intersection (noise-corrected squared distinctiveness rises ~2.5× from one to three features).
  • Systematic patterns
    • Model hierarchies of which dimensions matter are wrong and compressed: LLMs tend to flatten steering magnitudes across dimensions, exaggerating some weak signals and often failing to amplify the strongest real drivers.
    • Race and religion—among the strongest real predictors in ATP—are disproportionately ignored in multi-feature conditioning; party and gender are over-retained relative to their true importance.
  • Robustness
    • The collapse survives: aggregate-count readouts, individual-sampling readouts, and log‑prob/logit readouts; alternate prompts, role-frames, and chain-of-thought instructions do not meaningfully fix composition.
    • Results persist under stricter ground-truth cell size filters (n ≥ 200) and across model families and prompt variants.
  • Diagnostic baselines and calibration
    • The paper uses three methodological safeguards that materially affect interpretation:
    • Human sampling-noise floor to avoid spurious findings from small-real-cell noise.
    • Null calibration for similarity metrics (a random unrelated dimension already attains median cosine ≈ 0.84), so absolute similarities are not interpretable without a baseline.
    • Split-sample confirmation across waves to ensure replication.

Data & Methods

  • Data
    • Ground truth: respondent-level microdata from 15 waves of the Pew American Trends Panel (U.S. English, closed-form items).
    • Demographic dimensions: seven dimensions (e.g., race, religion, party, gender, age, income, education), yielding one-, two-, and three-way intersectional cells with at least n ≥ 20 respondents (larger-n subsets used for robustness).
  • Measurement constructs
    • Steering vector (sg): sg = pg − ppop (how subgroup answer distribution deviates from population).
    • Bias vector (eg): eg = ˆpg − pg (difference between modelled and real subgroup answers).
    • Pair-composition contest: for a two-feature pair AB, compare eAB (realized pair bias) against two hypotheses:
      • Additive hypothesis: eA + eB (sum of single-feature biases).
      • Collapse hypothesis: the best single-feature bias (either eA or eB).
    • Scoring by cosine similarity; wins are calibrated using synthetic additive-truth and collapse-truth cells built from measured single-feature biases plus empirically derived noise.
  • Calibration and noise controls
    • Synthetic endpoints give baseline win rates when truth is additive (best-single still wins ~40.3% due to noise/post-hoc max) and when truth is pure collapse (best-single wins ~84.3%). Observed rates are rescaled between these endpoints to form the collapse index.
    • Null baselines for cosine similarity computed (random unrelated dimension median cosine ≈ 0.84).
    • Split-sample and cluster-bootstrap procedures used to produce confidence intervals and avoid overfitting to specific waves.
  • Model and elicitation variants
    • Eight models spanning 7B open-weight families to frontier models (examples include GPT-4o-mini, GPT-5.5, Claude variants, Gemma-2-9B).
    • Readout paradigms: aggregate histogram of 1,000 simulated respondents; repeated single-roleplay draws (≈100 personas per cell); logit-based probability readout (OpinionQA style).
    • Prompting variations: neutral/system framings, chain-of-thought, explicit instructions to integrate identities—none materially changed collapse behavior.

Implications for AI Economics

  • Synthetic respondents are unreliable for intersectional inference
    • Economists should not assume LLMs can substitute for probability sampling of small intersectional subgroups. Synthetic surveys will likely under-represent joint identity effects and misattribute or ignore key drivers (notably race and religion).
    • Policy analysis, welfare estimates, or distributional predictions that depend on interaction effects across demographics (e.g., minority subgroups with particular political affiliations) risk systematic mismeasurement if based on LLM-generated synthetic samples.
  • Market research and demand estimation
    • Using LLMs to cheaply “create” rare consumer personas (e.g., Black conservative women in a local market) will likely yield outputs dominated by a single salient identity, biasing heterogeneity and interaction estimates used in segmentation, targeting, or willingness-to-pay studies.
  • Fairness, bias audits, and regulatory impact assessments
    • Audits that rely on silicon sampling to probe intersectional harms will understate intersectional effects and may miss the dimensions where harms concentrate (race, religion). Regulators and practitioners should treat silicon-sampling audits as incomplete unless rigorously validated against ground truth.
  • Econometric practice and model-based simulations
    • LLM-based simulation can be informative for marginal (single-dimension) analyses only with caution. For interaction effects, researchers should prefer real cross-tabulated data, hierarchical models (e.g., MAIHDA), or explicitly estimated interaction terms calibrated on real samples.
  • Practical recommendations for practitioners and researchers
    • Do not deploy LLM synthetic respondents as a substitute for real intersectional survey data without validation against joint empirical distributions.
    • If LLMs are used for exploratory work, report their limitations: validate which dimension the model is actually using, show collapse indices, and quantify uncertainty relative to noise baselines.
    • Adopt the paper’s suggested audit standards when evaluating silicon sampling: (1) control for human sampling noise; (2) calibrate similarity metrics against nulls and synthetic endpoints; (3) confirm findings on split samples/waves.
    • Where real intersectional data are unavailable, invest in targeted probability sampling or in statistical modeling approaches that combine limited real data with structural assumptions rather than relying purely on LLM outputs.
  • Broader economic inference caution
    • Because the collapse is model- and prompt-invariant and seen across architectures, the issue is structural to present LLM conditioning behavior. Economic conclusions that depend critically on multi-way heterogeneity should not rest on unvalidated LLM simulations.

If you want, I can: - Extract a short checklist you can use when considering LLM-based silicon sampling for a project. - Produce a one-page slide summarizing the paper’s figures and most important quantitative benchmarks for a presentation.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large-scale, multi-model empirical evaluation with explicit calibration and split-sample validation gives credible evidence about model behavior on the tested task, but findings are limited to one survey instrument (Pew ATP), US English closed-form items, a defined set of demographic dimensions, and the specific LLMs/configurations tested. Methods Rigorhigh — Careful benchmarking against respondent-level microdata, explicit noise-floor and null calibrations, split-sample replication, multiple elicitation paradigms, and evaluation across eight model families and many waves all increase internal validity and robustness; however, the study is observational with respect to model internals and limited to particular questions/modalities. SampleSimulated outputs from eight LLMs (including GPT-4o-mini, GPT-5.5, Claude variants, Gemma, etc.), producing full response distributions for 1,000 hypothetical respondents per demographic cell across one-, two-, and three-feature persona profiles drawn from seven demographic dimensions; scored against real respondent-level microdata from 15 waves of Pew Research Center's American Trends Panel (ATP) (2017–2021), using all intersectional cells with n ≥ 20 (and robustness checks at n ≥ 100, 200); total ~15.7 million simulated distributions in primary pipeline and ~21.1 million across conditions, evaluated over many closed-form ATP questions and across multiple elicitation paradigms (aggregate histogram, individual sampling, log-probability readout). Themeshuman_ai_collab adoption GeneralizabilityResults are limited to US English closed-form ATP survey items and may not hold for open-ended responses or different instruments., Findings apply to the specific LLMs, versions, and prompting variants tested; other models, fine-tuning, or system prompts could differ., Temporal limitation: ATP waves cover 2017–2021; population attitudes and model training data/post-training updates outside this window may alter results., Demographic coverage limited to the seven dimensions used; other identity axes or richer identity encodings (e.g., continuous attributes) were not tested., Does not evaluate downstream econometric use-cases directly (e.g., causal inference relying on synthetic respondents), so external validity for applied economic inference is uncertain.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Across 15 waves of the Pew American Trends Panel, the study evaluated synthetic respondents from eight language models against every real intersectional subgroup with at least 20 respondents. Output Quality other Accuracy and compositional validity of synthetic survey-response distributions
Reading fidelity high
Study strength high
n=8
0.3
Real intersectional subgroups are approximately additive in the direction and magnitude of their opinion shifts relative to the population. Output Quality positive Compositionality of subgroup opinion distributions
Reading fidelity high
Study strength high
n=236752
44.4% of pair-cell contests favored the best single feature, compared with a 49.3% noise-calibrated ceiling for perfect additivity
0.3
Two-feature personas produced by the language models are better explained by one retained feature than by the additive combination of both features in most cases. Output Quality negative Intersectional identity composition in simulated response distributions
Reading fidelity high
Study strength high
n=8
75.3–81.2% of profile–question cells
0.3
After calibration against known additive and collapse-generating synthetic cells, the models' two-feature conditioning shows a strong collapse toward one identity rather than composition of both. Output Quality negative Degree of collapse of multi-feature personas onto a single identity
Reading fidelity high
Study strength high
n=8
Self-calibrated collapse index 0.83–0.95
0.3
Adding a third demographic feature contributes very little additional compositional information in the simulated respondents. Output Quality negative Incremental contribution of a third identity feature to simulated responses
Reading fidelity high
Study strength medium
7.8–8.6% additive share
0.18
The observed identity collapse is robust to the response elicitation format, including aggregate distributions, individual persona sampling, and log-probability readouts. Output Quality negative Robustness of multi-feature identity collapse across elicitation methods
Reading fidelity high
Study strength medium
n=5
75.3–83.4% best-single win rate across model-by-paradigm cells
0.18
Prompting strategies, including explicit instructions to integrate both identities and chain-of-thought reasoning, do not materially eliminate the collapse. Output Quality negative Effect of prompting strategy on intersectional composition
Reading fidelity high
Study strength medium
≤2 percentage-point change for the reported framing variants
0.18
The feature retained by a model in a collapsed two-feature persona is only weakly related to which feature is actually most influential among real respondents. Decision Quality negative Relevance-tracking accuracy in selecting the dominant identity feature
Reading fidelity high
Study strength high
n=8
53.3–57.9% agreement, only 3–7 percentage points above permutation chance
0.3
Models systematically underrepresent race and religion, which are the strongest real drivers of opinion, while exaggerating the influence of weaker dimensions such as gender and party. Output Quality mixed Magnitude of demographic effects on simulated opinion distributions
Reading fidelity high
Study strength high
n=429
Human-to-model mean TV: gender 0.039→0.074; party 0.071→0.130; race 0.080→0.080; religion 0.088→0.093
0.3
Real subgroups become approximately 2.5 times more distinctive in squared distance from the population when moving from one to three intersecting features. Output Quality positive Distinctiveness of intersectional subgroup opinion distributions
Reading fidelity high
Study strength medium
2.5× increase in squared distinctiveness
0.18
The magnitude of real pair-subgroup opinion shifts is essentially additive rather than super-additive. Output Quality null_result Excess distinctiveness of pair intersections beyond additive single-feature effects
Reading fidelity high
Study strength high
−0.0008; 95% CI [−0.0011, −0.0006]
0.3

Notes