0 cumulative citations
View corpus contextLarge language models cannot reliably simulate intersectional survey respondents: two-feature personas usually reduce to one dominant identity and often ignore race and religion, meaning synthetic samples misrepresent intersectional opinion structure.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.
Summary
Main Finding
Large language models do not reliably simulate intersectional demographic identities. When prompted to answer as members of two- or three-feature demographic subgroups, LLMs typically “collapse” the persona onto one (or at most two) identity dimensions rather than integrating the multiple identities additively as real respondents do. This failure is large, robust across eight models and multiple prompting/readout paradigms, and systematically downweights race and religion—the strongest real drivers of opinion.
Key Points
- Scope and scale
- Evaluation against every real intersectional subgroup (n ≥ 20) in 15 waves of Pew’s American Trends Panel (ATP).
- 15.7 million simulated response distributions in the primary pipeline; 21.1 million across all conditions; eight LLMs tested (7B open models through current frontier models).
- Core quantitative results
- For two-feature personas, a single feature’s simulated bias explains the model’s pair output better than the additive combination in 75–82% of subgroups (raw win rates across models).
- Calibrated “collapse index” (rescaling wins between additive-truth and collapse-truth endpoints) places models about 0.83–0.95 of the way to pure collapse (e.g., GPT-4o-mini ≈ 0.86–0.87).
- Humans: pair identities are approximately additive. Real subgroups' weights center near full weight on both features (median (α, β) ≈ (0.95, 0.95), total α+β ≈ 1.83).
- Models: pair biases concentrate on a one-identity budget (median α+β ≈ 0.98), typically dominated by one feature (median dominant share ≈ 0.88).
- Real intersections grow more distinctive with intersection (noise-corrected squared distinctiveness rises ~2.5× from one to three features).
- Systematic patterns
- Model hierarchies of which dimensions matter are wrong and compressed: LLMs tend to flatten steering magnitudes across dimensions, exaggerating some weak signals and often failing to amplify the strongest real drivers.
- Race and religion—among the strongest real predictors in ATP—are disproportionately ignored in multi-feature conditioning; party and gender are over-retained relative to their true importance.
- Robustness
- The collapse survives: aggregate-count readouts, individual-sampling readouts, and log‑prob/logit readouts; alternate prompts, role-frames, and chain-of-thought instructions do not meaningfully fix composition.
- Results persist under stricter ground-truth cell size filters (n ≥ 200) and across model families and prompt variants.
- Diagnostic baselines and calibration
- The paper uses three methodological safeguards that materially affect interpretation:
- Human sampling-noise floor to avoid spurious findings from small-real-cell noise.
- Null calibration for similarity metrics (a random unrelated dimension already attains median cosine ≈ 0.84), so absolute similarities are not interpretable without a baseline.
- Split-sample confirmation across waves to ensure replication.
Data & Methods
- Data
- Ground truth: respondent-level microdata from 15 waves of the Pew American Trends Panel (U.S. English, closed-form items).
- Demographic dimensions: seven dimensions (e.g., race, religion, party, gender, age, income, education), yielding one-, two-, and three-way intersectional cells with at least n ≥ 20 respondents (larger-n subsets used for robustness).
- Measurement constructs
- Steering vector (sg): sg = pg − ppop (how subgroup answer distribution deviates from population).
- Bias vector (eg): eg = ˆpg − pg (difference between modelled and real subgroup answers).
- Pair-composition contest: for a two-feature pair AB, compare eAB (realized pair bias) against two hypotheses:
- Additive hypothesis: eA + eB (sum of single-feature biases).
- Collapse hypothesis: the best single-feature bias (either eA or eB).
- Scoring by cosine similarity; wins are calibrated using synthetic additive-truth and collapse-truth cells built from measured single-feature biases plus empirically derived noise.
- Calibration and noise controls
- Synthetic endpoints give baseline win rates when truth is additive (best-single still wins ~40.3% due to noise/post-hoc max) and when truth is pure collapse (best-single wins ~84.3%). Observed rates are rescaled between these endpoints to form the collapse index.
- Null baselines for cosine similarity computed (random unrelated dimension median cosine ≈ 0.84).
- Split-sample and cluster-bootstrap procedures used to produce confidence intervals and avoid overfitting to specific waves.
- Model and elicitation variants
- Eight models spanning 7B open-weight families to frontier models (examples include GPT-4o-mini, GPT-5.5, Claude variants, Gemma-2-9B).
- Readout paradigms: aggregate histogram of 1,000 simulated respondents; repeated single-roleplay draws (≈100 personas per cell); logit-based probability readout (OpinionQA style).
- Prompting variations: neutral/system framings, chain-of-thought, explicit instructions to integrate identities—none materially changed collapse behavior.
Implications for AI Economics
- Synthetic respondents are unreliable for intersectional inference
- Economists should not assume LLMs can substitute for probability sampling of small intersectional subgroups. Synthetic surveys will likely under-represent joint identity effects and misattribute or ignore key drivers (notably race and religion).
- Policy analysis, welfare estimates, or distributional predictions that depend on interaction effects across demographics (e.g., minority subgroups with particular political affiliations) risk systematic mismeasurement if based on LLM-generated synthetic samples.
- Market research and demand estimation
- Using LLMs to cheaply “create” rare consumer personas (e.g., Black conservative women in a local market) will likely yield outputs dominated by a single salient identity, biasing heterogeneity and interaction estimates used in segmentation, targeting, or willingness-to-pay studies.
- Fairness, bias audits, and regulatory impact assessments
- Audits that rely on silicon sampling to probe intersectional harms will understate intersectional effects and may miss the dimensions where harms concentrate (race, religion). Regulators and practitioners should treat silicon-sampling audits as incomplete unless rigorously validated against ground truth.
- Econometric practice and model-based simulations
- LLM-based simulation can be informative for marginal (single-dimension) analyses only with caution. For interaction effects, researchers should prefer real cross-tabulated data, hierarchical models (e.g., MAIHDA), or explicitly estimated interaction terms calibrated on real samples.
- Practical recommendations for practitioners and researchers
- Do not deploy LLM synthetic respondents as a substitute for real intersectional survey data without validation against joint empirical distributions.
- If LLMs are used for exploratory work, report their limitations: validate which dimension the model is actually using, show collapse indices, and quantify uncertainty relative to noise baselines.
- Adopt the paper’s suggested audit standards when evaluating silicon sampling: (1) control for human sampling noise; (2) calibrate similarity metrics against nulls and synthetic endpoints; (3) confirm findings on split samples/waves.
- Where real intersectional data are unavailable, invest in targeted probability sampling or in statistical modeling approaches that combine limited real data with structural assumptions rather than relying purely on LLM outputs.
- Broader economic inference caution
- Because the collapse is model- and prompt-invariant and seen across architectures, the issue is structural to present LLM conditioning behavior. Economic conclusions that depend critically on multi-way heterogeneity should not rest on unvalidated LLM simulations.
If you want, I can: - Extract a short checklist you can use when considering LLM-based silicon sampling for a project. - Produce a one-page slide summarizing the paper’s figures and most important quantitative benchmarks for a presentation.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Across 15 waves of the Pew American Trends Panel, the study evaluated synthetic respondents from eight language models against every real intersectional subgroup with at least 20 respondents. Output Quality | other | Accuracy and compositional validity of synthetic survey-response distributions |
Reading fidelity
high
Study strength
high
|
n=8
|
| Real intersectional subgroups are approximately additive in the direction and magnitude of their opinion shifts relative to the population. Output Quality | positive | Compositionality of subgroup opinion distributions |
Reading fidelity
high
Study strength
high
|
n=236752
44.4% of pair-cell contests favored the best single feature, compared with a 49.3% noise-calibrated ceiling for perfect additivity
|
| Two-feature personas produced by the language models are better explained by one retained feature than by the additive combination of both features in most cases. Output Quality | negative | Intersectional identity composition in simulated response distributions |
Reading fidelity
high
Study strength
high
|
n=8
75.3–81.2% of profile–question cells
|
| After calibration against known additive and collapse-generating synthetic cells, the models' two-feature conditioning shows a strong collapse toward one identity rather than composition of both. Output Quality | negative | Degree of collapse of multi-feature personas onto a single identity |
Reading fidelity
high
Study strength
high
|
n=8
Self-calibrated collapse index 0.83–0.95
|
| Adding a third demographic feature contributes very little additional compositional information in the simulated respondents. Output Quality | negative | Incremental contribution of a third identity feature to simulated responses |
Reading fidelity
high
Study strength
medium
|
7.8–8.6% additive share
|
| The observed identity collapse is robust to the response elicitation format, including aggregate distributions, individual persona sampling, and log-probability readouts. Output Quality | negative | Robustness of multi-feature identity collapse across elicitation methods |
Reading fidelity
high
Study strength
medium
|
n=5
75.3–83.4% best-single win rate across model-by-paradigm cells
|
| Prompting strategies, including explicit instructions to integrate both identities and chain-of-thought reasoning, do not materially eliminate the collapse. Output Quality | negative | Effect of prompting strategy on intersectional composition |
Reading fidelity
high
Study strength
medium
|
≤2 percentage-point change for the reported framing variants
|
| The feature retained by a model in a collapsed two-feature persona is only weakly related to which feature is actually most influential among real respondents. Decision Quality | negative | Relevance-tracking accuracy in selecting the dominant identity feature |
Reading fidelity
high
Study strength
high
|
n=8
53.3–57.9% agreement, only 3–7 percentage points above permutation chance
|
| Models systematically underrepresent race and religion, which are the strongest real drivers of opinion, while exaggerating the influence of weaker dimensions such as gender and party. Output Quality | mixed | Magnitude of demographic effects on simulated opinion distributions |
Reading fidelity
high
Study strength
high
|
n=429
Human-to-model mean TV: gender 0.039→0.074; party 0.071→0.130; race 0.080→0.080; religion 0.088→0.093
|
| Real subgroups become approximately 2.5 times more distinctive in squared distance from the population when moving from one to three intersecting features. Output Quality | positive | Distinctiveness of intersectional subgroup opinion distributions |
Reading fidelity
high
Study strength
medium
|
2.5× increase in squared distinctiveness
|
| The magnitude of real pair-subgroup opinion shifts is essentially additive rather than super-additive. Output Quality | null_result | Excess distinctiveness of pair intersections beyond additive single-feature effects |
Reading fidelity
high
Study strength
high
|
−0.0008; 95% CI [−0.0011, −0.0006]
|