22 cumulative citations
View corpus contextLarge language models can speed and cheapen social experiments, but only statistical calibration—not prompt tweaks—offers formal guarantees for causal inference when combined with auxiliary human data; heuristics help exploration but lack confirmation-level validity.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations support valid inference about human behavior. We contrast two strategies for obtaining valid estimates of causal effects and clarify the assumptions under which each is suitable for exploratory versus confirmatory research. Heuristic approaches seek to establish that simulated and observed human behavior are interchangeable through prompt engineering, model fine-tuning, and other repair strategies designed to reduce LLM-induced inaccuracies. While useful for many exploratory tasks, heuristic approaches lack the formal statistical guarantees typically required for confirmatory research. In contrast, statistical calibration combines auxiliary human data with statistical adjustments to account for discrepancies between observed and simulated responses. Under explicit assumptions, statistical calibration preserves validity and provides more precise estimates of causal effects at lower cost than experiments that rely solely on human participants. Yet the potential of both approaches depends on how well LLMs approximate the relevant populations. We consider what opportunities are overlooked when researchers focus myopically on substituting LLMs for human participants in a study.
Summary
Main Finding
The paper distinguishes two distinct, complementary paths for using large language models (LLMs) as substitutes or supplements for human subjects in behavioral research and evaluates when each can provide valid evidence. Heuristic (validate-then-simulate or simulate-then-validate) approaches are useful for exploratory work and design checks but lack formal statistical guarantees and can mask systematic biases. Statistical calibration—combining a (usually small) human sample with a larger set of LLM-generated responses and explicit statistical adjustments—can, under explicit assumptions, yield unbiased and more precise estimates at lower cost than human-only studies. The value of either approach depends critically on how well the LLM approximates the target human population and on whether calibration assumptions hold.
Key Points
- Two broad strategies:
- Heuristic validation: Demonstrate resemblance between LLM and human responses in some setting (direction/significance of effects, correlations, distributional distances, predictive accuracy, indistinguishability/Turing tests, face validity, expert appraisal, representational alignment) and then generalize to related scenarios. Best for exploratory work and instrument/design debugging.
- Statistical calibration: Use auxiliary human data plus statistical adjustments to correct discrepancies between LLM predictions and human outcomes; provides formal unbiasedness guarantees for target parameters under explicit assumptions.
- Empirical patterns in the literature:
- Moderate-to-high correlations between human and LLM effect sizes reported (e.g., ~0.5–0.85 across different study sets), but LLMs frequently overestimate effect magnitudes and sometimes produce significant effects not found in human samples.
- LLMs can reproduce directions of many effects but often differ in moments (means, variances) and can distort magnitudes or heterogeneity.
- Threats to heuristic substitution:
- Systematic bias/distortion (consistent over- or under-estimation; different variance structure).
- Memorization or training-data leakage that can masquerade as correct behavior.
- Lack of guarantees about generalization to novel prompts, populations, or counterfactual scenarios.
- Advantages and caveats of calibration:
- When calibration assumptions hold, augmenting a human sample with LLM data can reduce variance and cost, enabling better-powered tests (useful for small effects, interactions, or many conditions).
- Gains depend on LLM bias: if bias is large or poorly modeled, calibration offers little or no benefit and may mislead.
- Calibration requires explicit, justifiable assumptions (e.g., stable mapping from LLM outputs to human outcomes conditional on observables).
- Opportunities beyond simple substitution:
- LLMs can accelerate hypothesis generation, instrument debugging, design selection, and exploration of scenarios that are hard or impossible to observe with humans (historical agents, rare subpopulations), but these uses require different validation criteria than confirmatory inference.
Data & Methods
- Approach of the paper: conceptual analysis and synthesis of emerging empirical work on LLM simulations in behavioral science, combined with a formal framing of the data-generating process (f*), observed data D = {(Xi, Yi)}, and LLM outputs ˆf(X).
- Key modeling elements:
- Distinguishes datasets with joint human+LLM labels (Dshared = {(Xi, Yi, ˆf(Xi))}) used for validation, and larger LLM-only datasets (DLLM = {(˜Xi, ˆf(˜Xi))}) used for downstream analysis or augmentation.
- Formal discussion of target parameters (means, regression coefficients) and inference goals (hypothesis tests, decision rules).
- Summarized empirical validation metrics used across studies:
- Direction and significance concordance of effects.
- Correlation and comparison of effect size magnitudes between human and LLM results.
- Predictive accuracy at the individual and study level (in-sample and out-of-sample held-out tests).
- Distributional distances (Wasserstein, KL divergence, total variation, Earth mover distance).
- Human indistinguishability/Turing-style tests and face-validity or expert appraisal.
- Representational alignment via probing LLM internal activations vs. human neural data (e.g., fMRI correlations).
- The paper synthesizes threats (systematic bias, overfitting/memorization, miscalibrated uncertainty) and formal conditions required for statistical calibration to achieve unbiased inference.
Implications for AI Economics
- Appropriate uses in economics:
- Exploratory tasks: hypothesis generation, robustness checks, instrument and survey design, simulating counterfactuals or historically unobservable agents, and stress-testing experimental protocols.
- Power augmentation: hybrid designs that combine modest human samples with many LLM-generated responses can improve precision for small effects, interaction terms, or factorial designs—if calibration assumptions are plausible and validated.
- Cost-effective sensitivity analyses for policy evaluation and mechanism exploration before committing to costly human data collection.
- Key cautions for economists:
- Do not substitute LLM outputs for human data in confirmatory causal claims without calibration and transparent assumptions. Heuristic similarity (directional agreement, indistinguishability) does not imply unbiased causal estimates.
- LLMs often conform to theoretical rationales more than humans; this can inflate conformity to economic models and yield overconfident conclusions if taken as human evidence.
- Beware of overestimating effect magnitudes, underestimating heterogeneity, or relying on model behavior shaped by training-data artifacts (memorization).
- Recommended practices:
- Use a hybrid workflow: collect a Dshared sample, quantify LLM bias relative to human responses, and apply statistical calibration (weighting, regression adjustment, or other measurement-error corrections) with pre-specified assumptions and sensitivity analyses.
- Pre-register validation/calibration strategies and the assumptions needed for extrapolation from LLMs to target populations.
- Report both LLM-only, human-only, and calibrated estimates to show robustness and the influence of LLM augmentation.
- Evaluate distributional fidelity, not just direction/significance; examine moments, heterogeneity, and tail behavior relevant to economic parameters.
- Consider model provenance and training-data overlap with the target domain (risk of memorization) and test out-of-sample generalization to new prompts and populations.
- Frontier opportunities for economic research:
- Formal development and empirical testing of calibration methods tailored to economic estimands (treatment effects, demand elasticities, welfare measures).
- Use LLMs to explore complex policy counterfactuals, generate candidate mechanisms, and prioritize costly field experiments.
- Integrate representational alignment and model-internal probes to connect behavioral models, neural data, and economic theory—while carefully validating the external relevance of such links.
Summary takeaway: LLMs are powerful tools for exploration, design, and potential cost-saving augmentation of human data in economics, but credible causal inference requires explicit statistical calibration and careful validation; heuristic substitution is insufficient for confirmatory claims.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. Research Productivity | positive | cost-effectiveness and speed of generating responses (qualitative) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| There is limited guidance on when simulations using LLMs support valid inference about human behavior. Research Productivity | negative | availability of methodological guidance for valid inference from LLM-based simulations |
Reading fidelity
high
Study strength
low
|
not reported
|
| Two strategies are contrasted for obtaining valid estimates of causal effects when using LLMs as synthetic participants: (1) heuristic approaches (prompt engineering, model fine-tuning, repair strategies) and (2) statistical calibration combining auxiliary human data with statistical adjustments. Research Productivity | mixed | methods for obtaining valid causal estimates from LLM-based simulations |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Heuristic approaches (prompt engineering, fine-tuning, repair strategies) are useful for many exploratory tasks but lack the formal statistical guarantees typically required for confirmatory research. Research Productivity | mixed | suitability of heuristic approaches for exploratory vs. confirmatory research (validity guarantees) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Statistical calibration, which combines auxiliary human data with statistical adjustments, preserves validity under explicit assumptions and can provide more precise estimates of causal effects at lower cost than experiments relying solely on human participants. Research Productivity | positive | validity and precision of causal effect estimates; cost relative to human-only experiments |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The potential of both heuristic and statistical calibration approaches depends on how well LLMs approximate the relevant populations. Research Productivity | neutral | degree to which LLMs approximate relevant human populations (affecting validity of methods) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Focusing narrowly on substituting LLMs for human participants overlooks other opportunities (i.e., there are research opportunities beyond direct substitution). Research Productivity | neutral | breadth of research opportunities when using LLMs (conceptual) |
Reading fidelity
high
Study strength
low
|
not reported
|