0 cumulative citations
View corpus contextLarge language models favour wealthier, tech-advanced countries when simulating public values, producing systematic cross-country representational inequality; common fixes such as native-language prompting, post-training and alignment often raise average accuracy but fail to produce consistent gains in equality.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Traditional methods for studying human opinions often struggle to support representative and scalable research across countries. Large language models (LLMs) can serve as scalable proxies for simulating human opinions, enabling more efficient opinion analysis. However, this use of LLMs requires not only high average accuracy but also representational equality, that is, comparable simulation accuracy across populations. Uneven simulation accuracy may reproduce or amplify societal biases in downstream applications. This study systematically investigates country-level representational equality across 59 countries and finds substantial, systematic inequality. Populations from wealthier and more technologically advanced countries are simulated more accurately. We further compare two foundational intervention pathways, contextual adaptation and parametric modification, and show that improvements in average or target-group accuracy do not necessarily translate into greater representational equality. For contextual adaptation, native-language prompting generally improves accuracy but remains model-dependent, whereas additional information more often improves both accuracy and equality. For parametric modification, language-specific continued post-training improves accuracy for targeted language groups but unevenly, while preference alignment yields no systematic gains in accuracy or equality. Human-annotated preference data generally preserve accuracy better than AI-annotated data. These findings highlight the need for representational equality alongside accuracy and offer guidance for more inclusive, socially responsible LLM-based simulations.
Summary
Main Finding
LLMs can approximate cross-country human value distributions but do so unevenly: simulation accuracy is systematically higher for populations in wealthier, more technologically advanced countries. Common intervention strategies (contextual adaptation and parametric modification) can raise average accuracy or improve performance for target groups, but these gains do not reliably translate into more equal (cross-country) simulation performance. Some interventions (e.g., adding contextual information; language-specific post-training) help accuracy but often increase inequality or do so unevenly; preference-alignment interventions show no systematic improvement in equality. Human-annotated preference data preserve accuracy better than AI-annotated data.
Key Points
- Representational Equality: The paper formalizes "representational equality" as the evenness of LLM simulation accuracy across populations (here, countries). It is a diagnostic objective distinct from overall accuracy.
- Scope: Evaluation across 59 countries and 2,420 demographic subpopulations (country × demographics: gender, age, education, income).
- Core result: Simulation accuracy is systematically higher for countries with greater economic and technological development (e.g., higher GDP per capita, internet use, Global Innovation Index).
- Primary equality metric: coefficient of variation (CV) of country-level accuracies; also reported max-min, min-max ratio, Gini; composite AE metric (geometric mean of average JSD and EqCV) combines accuracy and equality.
- Output extraction: first-token probability method over answer-option tokens (top-20 tokens, capped residual mass for missing options). Validated against text generation approach.
- Simulation accuracy measure: 1 − Jensen–Shannon divergence between model-predicted and empirical (World Values Survey) response distributions; question-level accuracies averaged by value dimension then across dimensions.
- Interventions evaluated:
- Contextual adaptation: native-language prompting (EN, ZH, AR, ES) and adding additional information (e.g., historical memory, persona profile, few-shot examples). Findings: native-language prompting usually improves accuracy but is model-dependent; providing additional information more often improves both accuracy and equality.
- Parametric modification: language-specific continued post-training and different alignment sources (human-annotated vs AI-annotated preference data). Findings: post-training on native-language data improves target-language accuracy but unevenly across countries/languages; alignment using preference data produced no systematic gains in equality; human-annotated alignment data better preserve accuracy than AI-annotated.
- Diagnostic correlates: Using a PEST framework (Political, Economic, Socio-cultural, Technological), strong associations found between country-level simulation accuracy and economic/technological indicators (GDP per capita, internet penetration, GII); governance and cultural dimensions also explored but economic/tech factors are prominent drivers.
- Practical takeaway: Improving average accuracy is not sufficient—models and mitigation strategies must be evaluated on representational equality to avoid reproducing or amplifying cross-country disparities.
Data & Methods
- Empirical target: World Values Survey (WVS) response distributions across countries and demographic subpopulations.
- Countries: 59 countries selected for coverage and comparability.
- Subpopulations: 2,420 cells combining country and demographic attributes (gender, age, education, income).
- Models: Multiple LLMs evaluated (various scales and families; exact model list in paper).
- Response extraction:
- First-token probability method: request top-20 first-token log probabilities, filter tokens matching choice labels (A, B, C...), normalize, and handle missing options via capped residual-mass approximation.
- Validated consistency vs. free-text extraction in Appendix C.1.
- Accuracy computation:
- Per question: Acc(H, M, q) = 1 − JSD(D_human(H,q), D_model(M,H,q))
- Aggregate across questions balanced by value dimension, average subpopulations into country-level accuracy Ac.
- Representational equality:
- Primary index: EqCV = σ(Ac) / μ(Ac) (coefficient of variation across countries).
- Additional indices: max-min, min-max ratio, Gini—reported to show convergence.
- Composite AE metric: sqrt(JSD_avg × EqCV) (lower is better).
- Interventions tested:
- Contextual: prompting in native languages (EN, ZH, AR, ES), adding persona/profile info, few-shot examples, historical memory.
- Parametric: continued post-training (language-specific corpora), alignment using preference datasets from human annotators vs. AI annotators; variants of alignment settings.
- Macro-level correlational analysis:
- PEST indicators: GDP per capita, Internet use, Global Innovation Index (GII), Worldwide Governance Indicators (WGI), Hofstede cultural dimensions.
- Correlation and regression-style diagnostics assessing associations between these indicators and country-level simulation accuracy.
Implications for AI Economics
- Distributional consequences of LLM-based tools:
- LLMs’ uneven simulation fidelity implies economic and policy analyses that rely on model-based population simulations risk misrepresenting preferences and behaviors in lower-income or less digitally visible countries, potentially biasing policy recommendations and allocation decisions.
- Market research, demand estimation, and forecasting that use LLM simulations may systematically overfit to perspectives of wealthier/connected populations unless equality is explicitly addressed.
- Investment and data priorities:
- Improving representational equality likely requires targeted investment in data collection and curation for underrepresented countries and languages (e.g., native-language corpora, survey augmentation). From an economic perspective, funders and firms aiming for global reach should budget for such targeted data efforts.
- Infrastructure and digital inclusion (internet access, digital literacy) indirectly affect LLM coverage; public and private investments that raise digital visibility can improve downstream model representativeness.
- Model evaluation and procurement:
- Procurement and regulatory standards for LLMs used in economic policy, international development, or cross-country market analysis should include representational equality metrics (e.g., EqCV, AE composite) alongside average accuracy.
- Contracts for model deployment could require transparency on cross-country performance and mitigation plans for low-coverage regions.
- Design of mitigation and product strategies:
- Contextual adaptation (e.g., native-language prompts and richer contextual profiles) is a low-cost, often effective approach but is model-dependent; firms should empirically test prompts per language/market rather than assume uniform gains.
- Parametric fixes (post-training, alignment) can improve performance for targeted languages but may exacerbate inequality unless applied across languages/countries proportionally—costs scale with the number of targeted locales.
- Human-annotated alignment data are preferable where preserving overall accuracy is important; economic trade-offs exist between the higher cost of human annotation and the quality/equity benefits.
- Policy and fairness regulation:
- Regulators and international organizations should consider representational equality when evaluating AI tools used in global policymaking, aid allocation, and international surveys; unequal simulation fidelity may produce biased policy signals.
- Funding and incentives (e.g., grants, procurement preferences) could be structured to reward models and providers that demonstrate better representational equality.
- Research and cost-effectiveness:
- For researchers using LLMs as scalable proxies for cross-country social science, the paper highlights a cost/accuracy/equality trade-off: cheaper methods (prompting) may yield improvements but are brittle; parametric methods are costlier and can be inequitably effective. Economic analyses should account for these trade-offs in study design and budgeting.
- Suggested practical steps for economists and organizations:
- Always report cross-country accuracy and equality metrics when using LLM-based simulations.
- Pilot native-language contextual prompts and additional contextual data, and validate results against empirical surveys where possible.
- Prioritize human-annotated alignment for high-stakes applications; if using AI-annotated data, validate downstream effects on both accuracy and equality.
- Consider targeted data collection investments for underrepresented countries as part of AI deployment budgets to reduce representational inequality.
Overall, the paper cautions that high average simulation performance hides systematic cross-country disparities with tangible economic and policy consequences. Addressing representational equality requires measured intervention, targeted data investments, and adoption of equality-aware evaluation and procurement practices.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLM-based value simulation exhibits substantial and systematic representational inequality across countries. Inequality | negative | Equality of simulation accuracy across country populations |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Populations from wealthier and more technologically advanced countries are simulated more accurately by the evaluated LLMs. Output Quality | positive | Similarity between model-simulated and empirical human value-response distributions |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Improving average simulation accuracy or the accuracy of a target group does not necessarily improve representational equality. Inequality | mixed | Average simulation accuracy and dispersion of accuracy across countries |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Native-language prompting generally improves simulation accuracy, but the effect depends on the model and does not consistently resolve representational inequality. Output Quality | mixed | Simulation accuracy and equality across language and country groups |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Providing additional information in prompts more often improves both simulation accuracy and representational equality. Output Quality | positive | Simulation accuracy and equality of accuracy across populations |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Language-specific continued post-training improves accuracy for targeted language groups, but the improvements are uneven across groups. Output Quality | mixed | Simulation accuracy among targeted language groups and its distribution across groups |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Preference alignment produces no systematic improvements in either simulation accuracy or representational equality. Inequality | null_result | Simulation accuracy and representational equality |
Reading fidelity
high
Study strength
medium
|
n=59
|
| Human-annotated preference data preserve simulation accuracy better than AI-annotated preference data. Output Quality | positive | Preservation of simulation accuracy after preference-based training |
Reading fidelity
high
Study strength
medium
|
n=59
|
| The study evaluates 2,420 demographic subpopulations across 59 countries for each model and intervention setting. Other | other | Coverage of evaluated demographic subpopulations |
Reading fidelity
high
Study strength
high
|
n=2420
|