Digital twins can mimic average human responses but rarely substitute for them: while LLMs reproduce many aggregate behavioral effects, they usually fail to predict which individuals drive those effects, so newer models and extra respondent data improve alignment but do not reliably reduce the need for human measurements.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.
Summary
Main Finding
The authors introduce "statistical substitutability"—an estimand-specific criterion that asks whether LLM-based digital twins can reduce human measurement while preserving valid inference—and show empirically that current digital twins often reproduce aggregate human effects but lack the respondent-level signal and stable calibration needed to reliably substitute for human data. Improvements in model capacity or richer respondent conditioning sometimes improve aggregate fidelity but do not consistently translate into trustworthy reductions in required human samples.
Key Points
- Definition: Statistical substitutability measures whether twin predictions can reduce human measurement for a prespecified human estimand while (a) preserving valid inference and (b) achieving at least the reference human-only precision.
- Four evaluation dimensions:
- Aggregate fidelity — do twin predictions reproduce population means/distributions/effects?
- Paired respondent-level signal — do twin predictions track which individuals deviate from group means (necessary for calibration gains)?
- Finite-sample human-label recovery — can limited human labels + predictions (via prediction-powered inference, PPI) recover human-only precision?
- Transport/stability — does calibration/generalization hold across populations or samples?
- Main empirical results:
- Across Twin-2K reconstructions and a Moore–Berg extension, twins frequently reproduce average human effects but provide little information about individual-level deviations that matter for improving inference.
- Newer models and richer respondent information can improve some fidelity and paired-signal metrics, but these improvements do not reliably convert into stable human-data savings.
- Human calibration (using labeled human samples) can reduce aggregate prediction error, but limited labeled samples often fail to produce stable precision gains or reliable transportability.
- Conceptual takeaway: Behavioral fidelity (matching means, distributions, or aggregate effects) is neither necessary nor sufficient for a digital twin to be useful as a substitute in confirmatory inference. The right evaluation target is inferential usefulness for the estimand, not only behavioral realism.
Data & Methods
- Framework:
- Builds on prediction-powered inference (PPI) and mixed-subject designs that combine model predictions with a smaller human validation sample to correct prediction error and potentially improve precision.
- Treats twin outputs as prediction-only data (not human observations) and evaluates whether PPI can recover the reference human estimand and precision using fewer human labels plus predictions.
- Empirical evaluations:
- Twin-2K dataset (Toubia et al., 2025): reconstruction of 12 behavioral studies using the Twin-2K digital twins to compare aggregate replication vs. respondent-level signal and finite-sample PPI performance.
- Moore–Berg extension: a partisan meta-perception study where the estimand is closely tied to participant characteristics (a favorable case for person-specific prediction). This extension tests multiple models, respondent representations, and calibration strategies, and examines transport when Twin-2K lacks the target outcomes.
- Models and comparisons:
- Multiple LLMs and respondent encodings (demographic/persona vs. information-rich histories).
- Calibration strategies include human-label calibration applied within-sample and transported across samples.
- Evaluation metrics:
- Aggregate alignment of means and treatment effects.
- Paired human–twin association for the estimand (correlation / signal that drives effective sample-size gains).
- PPI-based finite-sample precision and whether reduced-human designs can match the full human benchmark.
- Stability/transport tests for carrying calibration across populations.
Implications for AI Economics
- Valuing AI for measurement reduction requires estimand-specific evidence, not headline predictive metrics. Economic assessments that assume cost savings from replacing human respondents with LLM outputs will be overly optimistic unless paired-signal and calibration stability are demonstrated.
- Effective sample-size and power calculations must incorporate the paired correlation between predictions and human outcomes. ROI and budget trade-offs depend on that correlation and on the cost of obtaining the calibration (human labels) needed to realize gains.
- Policy and procurement: organizations should require inferential validation (statistical substitutability tests) for AI tools proposed to replace human data collection. Contracts that pay for "accuracy" alone are insufficient; buyers need guarantees about inferential validity and transport stability.
- Experimental design guidance:
- When considering mixed human–AI data, plan for a labeled validation subsample representative of the target estimand and population; allocate budget to obtain it.
- Use PPI-style calibrated estimators rather than naive substitution of model outputs as if they were ground-truth human data.
- Test transport assumptions before applying calibration across populations; unstable transport undermines claimed human-data savings.
- Research economics: the marginal value of additional model capacity or richer conditioning data should be judged by how much they increase paired respondent-level signal and stable calibration — not merely by improved aggregate accuracy. Investments in models or respondent data are worthwhile only if they reliably reduce human-sample requirements for target estimands.
- Limitations and research needs: current evidence is limited to behavioral experimental tasks (Twin-2K and Moore–Berg). For broader economic conclusions, further work is needed across domains, outcome types, longer-term stability, and deployment settings (e.g., market forecasting, consumer choice, high-stakes measurement).
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Across two empirical evaluations, LLM-based digital twins reproduced average human effects while providing little information about which individuals differed from those averages. Output Quality | mixed | Aggregate behavioral-effect replication and respondent-level predictive signal |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Newer models and richer respondent information improved some performance dimensions but did not reliably translate into savings in human data collection. Task Allocation | mixed | Reduction in required human measurement and related substitutability dimensions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Human calibration can reduce aggregate prediction error, but limited labeled samples often fail to produce stable precision gains. Decision Quality | mixed | Aggregate prediction error and precision of prediction-assisted inference |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Behavioral fidelity is neither necessary nor sufficient for statistical substitutability. Decision Quality | negative | Ability of digital-twin predictions to reduce human measurement while preserving valid inference |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper argues that digital twins should be judged for confirmatory use by whether they reduce uncertainty about human quantities, rather than merely by whether they reproduce human means, distributions, or effects. Decision Quality | positive | Uncertainty and inferential precision for human estimands |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In the Twin-2K-500 benchmark summarized by the paper, baseline twins achieved 71.7% average predictive accuracy, approximately 87.7% of the human test-retest benchmark, and reproduced about half of the evaluated behavioral effects. Output Quality | mixed | Held-out response predictive accuracy and behavioral-effect replication |
Reading fidelity
high
Study strength
medium
|
71.7% average predictive accuracy; 87.7% of the human test-retest benchmark; about half of evaluated behavioral effects
|
| Aggregate agreement between twin predictions and human outcomes does not establish that the predictions contain the respondent-level signal needed for human-data savings. Task Allocation | negative | Paired respondent-level information available for reducing human measurement |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Human validation data can correct average prediction error but cannot create respondent-level signal that is absent from the digital-twin predictions. Decision Quality | negative | Respondent-level predictive signal and potential precision gains from calibration |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Statistical substitutability is supported only when a reduced-human design preserves valid inference for the target estimand and attains at least the precision of the full human reference design. Decision Quality | positive | Validity and precision of inference for a prespecified human estimand |
Reading fidelity
high
Study strength
high
|
not reported
|