The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models can speed and cheapen social experiments, but only statistical calibration—not prompt tweaks—offers formal guarantees for causal inference when combined with auxiliary human data; heuristics help exploration but lack confirmation-level validity.

This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw · February 17, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jessica Hullman unresolved corpus identity
  2. David Broska unresolved corpus identity
  3. Huaman Sun unresolved corpus identity
  4. Aaron Shaw unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jessica Hullman provider ID
  2. David Broska provider ID
  3. Huaman Sun provider ID
  4. Aaron Shaw provider ID
The paper contrasts heuristic prompt-based repairs and formal statistical calibration for using LLMs as synthetic experiment participants, arguing that calibration—when its explicit assumptions hold—can produce valid and more precise causal estimates while heuristics are mainly suitable for exploratory work.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations support valid inference about human behavior. We contrast two strategies for obtaining valid estimates of causal effects and clarify the assumptions under which each is suitable for exploratory versus confirmatory research. Heuristic approaches seek to establish that simulated and observed human behavior are interchangeable through prompt engineering, model fine-tuning, and other repair strategies designed to reduce LLM-induced inaccuracies. While useful for many exploratory tasks, heuristic approaches lack the formal statistical guarantees typically required for confirmatory research. In contrast, statistical calibration combines auxiliary human data with statistical adjustments to account for discrepancies between observed and simulated responses. Under explicit assumptions, statistical calibration preserves validity and provides more precise estimates of causal effects at lower cost than experiments that rely solely on human participants. Yet the potential of both approaches depends on how well LLMs approximate the relevant populations. We consider what opportunities are overlooked when researchers focus myopically on substituting LLMs for human participants in a study.

Summary

Main Finding

The paper distinguishes two distinct, complementary paths for using large language models (LLMs) as substitutes or supplements for human subjects in behavioral research and evaluates when each can provide valid evidence. Heuristic (validate-then-simulate or simulate-then-validate) approaches are useful for exploratory work and design checks but lack formal statistical guarantees and can mask systematic biases. Statistical calibration—combining a (usually small) human sample with a larger set of LLM-generated responses and explicit statistical adjustments—can, under explicit assumptions, yield unbiased and more precise estimates at lower cost than human-only studies. The value of either approach depends critically on how well the LLM approximates the target human population and on whether calibration assumptions hold.

Key Points

  • Two broad strategies:
    • Heuristic validation: Demonstrate resemblance between LLM and human responses in some setting (direction/significance of effects, correlations, distributional distances, predictive accuracy, indistinguishability/Turing tests, face validity, expert appraisal, representational alignment) and then generalize to related scenarios. Best for exploratory work and instrument/design debugging.
    • Statistical calibration: Use auxiliary human data plus statistical adjustments to correct discrepancies between LLM predictions and human outcomes; provides formal unbiasedness guarantees for target parameters under explicit assumptions.
  • Empirical patterns in the literature:
    • Moderate-to-high correlations between human and LLM effect sizes reported (e.g., ~0.5–0.85 across different study sets), but LLMs frequently overestimate effect magnitudes and sometimes produce significant effects not found in human samples.
    • LLMs can reproduce directions of many effects but often differ in moments (means, variances) and can distort magnitudes or heterogeneity.
  • Threats to heuristic substitution:
    • Systematic bias/distortion (consistent over- or under-estimation; different variance structure).
    • Memorization or training-data leakage that can masquerade as correct behavior.
    • Lack of guarantees about generalization to novel prompts, populations, or counterfactual scenarios.
  • Advantages and caveats of calibration:
    • When calibration assumptions hold, augmenting a human sample with LLM data can reduce variance and cost, enabling better-powered tests (useful for small effects, interactions, or many conditions).
    • Gains depend on LLM bias: if bias is large or poorly modeled, calibration offers little or no benefit and may mislead.
    • Calibration requires explicit, justifiable assumptions (e.g., stable mapping from LLM outputs to human outcomes conditional on observables).
  • Opportunities beyond simple substitution:
    • LLMs can accelerate hypothesis generation, instrument debugging, design selection, and exploration of scenarios that are hard or impossible to observe with humans (historical agents, rare subpopulations), but these uses require different validation criteria than confirmatory inference.

Data & Methods

  • Approach of the paper: conceptual analysis and synthesis of emerging empirical work on LLM simulations in behavioral science, combined with a formal framing of the data-generating process (f*), observed data D = {(Xi, Yi)}, and LLM outputs ˆf(X).
  • Key modeling elements:
    • Distinguishes datasets with joint human+LLM labels (Dshared = {(Xi, Yi, ˆf(Xi))}) used for validation, and larger LLM-only datasets (DLLM = {(˜Xi, ˆf(˜Xi))}) used for downstream analysis or augmentation.
    • Formal discussion of target parameters (means, regression coefficients) and inference goals (hypothesis tests, decision rules).
  • Summarized empirical validation metrics used across studies:
    • Direction and significance concordance of effects.
    • Correlation and comparison of effect size magnitudes between human and LLM results.
    • Predictive accuracy at the individual and study level (in-sample and out-of-sample held-out tests).
    • Distributional distances (Wasserstein, KL divergence, total variation, Earth mover distance).
    • Human indistinguishability/Turing-style tests and face-validity or expert appraisal.
    • Representational alignment via probing LLM internal activations vs. human neural data (e.g., fMRI correlations).
  • The paper synthesizes threats (systematic bias, overfitting/memorization, miscalibrated uncertainty) and formal conditions required for statistical calibration to achieve unbiased inference.

Implications for AI Economics

  • Appropriate uses in economics:
    • Exploratory tasks: hypothesis generation, robustness checks, instrument and survey design, simulating counterfactuals or historically unobservable agents, and stress-testing experimental protocols.
    • Power augmentation: hybrid designs that combine modest human samples with many LLM-generated responses can improve precision for small effects, interaction terms, or factorial designs—if calibration assumptions are plausible and validated.
    • Cost-effective sensitivity analyses for policy evaluation and mechanism exploration before committing to costly human data collection.
  • Key cautions for economists:
    • Do not substitute LLM outputs for human data in confirmatory causal claims without calibration and transparent assumptions. Heuristic similarity (directional agreement, indistinguishability) does not imply unbiased causal estimates.
    • LLMs often conform to theoretical rationales more than humans; this can inflate conformity to economic models and yield overconfident conclusions if taken as human evidence.
    • Beware of overestimating effect magnitudes, underestimating heterogeneity, or relying on model behavior shaped by training-data artifacts (memorization).
  • Recommended practices:
    • Use a hybrid workflow: collect a Dshared sample, quantify LLM bias relative to human responses, and apply statistical calibration (weighting, regression adjustment, or other measurement-error corrections) with pre-specified assumptions and sensitivity analyses.
    • Pre-register validation/calibration strategies and the assumptions needed for extrapolation from LLMs to target populations.
    • Report both LLM-only, human-only, and calibrated estimates to show robustness and the influence of LLM augmentation.
    • Evaluate distributional fidelity, not just direction/significance; examine moments, heterogeneity, and tail behavior relevant to economic parameters.
    • Consider model provenance and training-data overlap with the target domain (risk of memorization) and test out-of-sample generalization to new prompts and populations.
  • Frontier opportunities for economic research:
    • Formal development and empirical testing of calibration methods tailored to economic estimands (treatment effects, demand elasticities, welfare measures).
    • Use LLMs to explore complex policy counterfactuals, generate candidate mechanisms, and prioritize costly field experiments.
    • Integrate representational alignment and model-internal probes to connect behavioral models, neural data, and economic theory—while carefully validating the external relevance of such links.

Summary takeaway: LLMs are powerful tools for exploration, design, and potential cost-saving augmentation of human data in economics, but credible causal inference requires explicit statistical calibration and careful validation; heuristic substitution is insufficient for confirmatory claims.

Assessment

Paper Typetheoretical Evidence Strengthn/a — Conceptual/methodological paper without original empirical tests; offers formal arguments and assumptions rather than empirical validation, so empirical evidence strength is not applicable. Methods Rigormedium — Provides a clear conceptual taxonomy and formal discussion of identification assumptions and calibration procedures, but lacks empirical validation, sensitivity analyses, or extensive simulation results in the paper itself to demonstrate practical performance across settings. SampleNo original human-subject data; synthesizes literature and hypothetical examples involving LLM-generated responses and auxiliary human datasets (e.g., small calibration samples, surveys, or experiment results) that would be combined with simulations for statistical adjustment. Themeshuman_ai_collab adoption IdentificationCompares two strategies: (1) heuristic interchangeability—attempt to make LLM responses substitutable for humans via prompt engineering, fine-tuning, and other repairs (no formal identification); (2) statistical calibration—use auxiliary human data to estimate the discrepancy between LLM-simulated and human responses and apply statistical adjustments (bias-correction, reweighting/transportability adjustments) to recover causal effects under explicit assumptions (e.g., transportability/exchangeability, correct specification of adjustment model, overlap between simulated and human covariate support). GeneralizabilityRelies on how well specific LLMs approximate the target human population—results may not generalize across models, tasks, or updates to LLMs., Statistical calibration requires auxiliary human data whose availability and quality vary by domain and population., Assumptions (transportability, correct adjustment model, overlap) may fail in many real-world settings, limiting validity., Heuristic repairs (prompt engineering, fine-tuning) are task- and context-specific and often do not transfer., Findings are methodological and may not straightforwardly translate into guidance for field experiments, high-stakes policy evaluation, or settings with complex incentives.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. Research Productivity positive cost-effectiveness and speed of generating responses (qualitative)
Reading fidelity high
Study strength speculative
not reported
0.02
There is limited guidance on when simulations using LLMs support valid inference about human behavior. Research Productivity negative availability of methodological guidance for valid inference from LLM-based simulations
Reading fidelity high
Study strength low
not reported
0.06
Two strategies are contrasted for obtaining valid estimates of causal effects when using LLMs as synthetic participants: (1) heuristic approaches (prompt engineering, model fine-tuning, repair strategies) and (2) statistical calibration combining auxiliary human data with statistical adjustments. Research Productivity mixed methods for obtaining valid causal estimates from LLM-based simulations
Reading fidelity high
Study strength speculative
not reported
0.02
Heuristic approaches (prompt engineering, fine-tuning, repair strategies) are useful for many exploratory tasks but lack the formal statistical guarantees typically required for confirmatory research. Research Productivity mixed suitability of heuristic approaches for exploratory vs. confirmatory research (validity guarantees)
Reading fidelity high
Study strength medium
not reported
0.12
Statistical calibration, which combines auxiliary human data with statistical adjustments, preserves validity under explicit assumptions and can provide more precise estimates of causal effects at lower cost than experiments relying solely on human participants. Research Productivity positive validity and precision of causal effect estimates; cost relative to human-only experiments
Reading fidelity high
Study strength medium
not reported
0.12
The potential of both heuristic and statistical calibration approaches depends on how well LLMs approximate the relevant populations. Research Productivity neutral degree to which LLMs approximate relevant human populations (affecting validity of methods)
Reading fidelity high
Study strength speculative
not reported
0.02
Focusing narrowly on substituting LLMs for human participants overlooks other opportunities (i.e., there are research opportunities beyond direct substitution). Research Productivity neutral breadth of research opportunities when using LLMs (conceptual)
Reading fidelity high
Study strength low
not reported
0.06

Notes