0 cumulative citations
View corpus contextGPT-3.5 mirrors UK household inflation perceptions at short horizons and across demographic groups, yet fails to internalize the post-2021 inflation surge because of its training cutoff and shows no consistent model of consumer-price dynamics.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.
Summary
Main Finding
GPT-3.5-turbo (gpt-3.5-turbo-0613), when prompted to act as households in a synthetic version of the Bank of England’s Inflation Attitudes Survey (IAS), can reproduce aggregate survey distributions and many demographic regularities in inflation perceptions and short-horizon expectations. A Shapley-value decomposition adapted to the synthetic-survey setting shows prompt-conditioned drivers of responses (notably a human-like oversensitivity to food inflation). However, the model lacks a coherent internal model of consumer-price inflation: micro-level correspondence with individual human responses is weak and partially unstable, and the LLM exhibits internal inconsistencies that limit its use for individual-level inference.
Key Points
- Quasi-experimental design: exploited GPT’s training cut-off (Sept 2021) to probe out-of-sample behaviour during the UK inflation surge (peaking late 2022).
- Synthetic surveys: GPT was prompted with synthetic personas (demographics) and economic information (subcomponent inflation) to answer IAS questions on perceived current inflation and expectations at 1, 2 and 5 years.
- Aggregate fit: with temperature tuning, GPT can match aggregate response distributions and short-horizon expectations fairly well.
- Meso-level alignment: GPT replicates key demographic correlations observed in real IAS responses (income, housing tenure, social class) and often aligns more closely with official statistics than with human survey responses.
- Component sensitivity: GPT shows heightened sensitivity to salient components (food/restaurant inflation), similar to documented human biases.
- Micro-level limits: individual-level responses are unstable across runs and inconsistent in ways that indicate GPT does not have a consistent world model of CPI dynamics.
- Explainability: introduced a Shapley-value decomposition for LLM outputs in the synthetic-survey setting to quantify contributions of prompt bits (demographics, each inflation component) and their interactions.
- Practical cautions: results depend on prompt design, temperature, and the specific model/version; ethical concerns and demographic biases persist. The model is used as a representative testbed, not an endorsement of any vendor.
Data & Methods
- Model and API: OpenAI chat completions API using gpt-3.5-turbo-0613. System + user prompts specified a persona (age, gender, region, income, education, tenure, social class) and economic conditioning.
- Economic conditioning: three-month averages of year-on-year inflation for CPIH subcomponents immediately preceding the survey date were given in prompts:
- Cross-validation sample (Nov 2022): food 15.0%, restaurants 8.1%, energy 76.0%, other 6.0%, CPIH 9.3%.
- Main test sample (Feb 2023): food 17.0%, restaurants 9.8%, energy 88.0%, other 5.0%, CPIH 9.2%.
- These values are well outside GPT’s training-range (pre-Oct 2021 averages: CPIH ≈ 2.6%).
- Experimental design:
- Two samples: 2022Q4 (Nov 2022) used for cross-validation/tuning; 2023Q1 (Feb 2023) as main test.
- Temperature grid: T ∈ {0, 0.25, 0.5, 0.75, 1, 1.25, 1.5} to control deterministic vs. stochastic outputs; temperature tuned on CV sample to match aggregate moments.
- Randomization: response option order was scrambled per respondent to avoid order bias.
- Treatments: presence vs omission (or baseline historic averages) of economic-condition bits in prompts defined treated vs control states; because the same synthetic subject can be re-run under different treatments, individual treatment effects were computed directly.
- Explainability: applied a Shapley-value decomposition (adapted for discrete prompt-treatment bits) to attribute contributions of each prompt element (e.g., food inflation, energy inflation, demographics) and their interactions to the model’s chosen survey response.
- Benchmarks: compared synthetic GPT responses at micro (individual), meso (demographic groups), and macro (aggregate distributions) levels against actual IAS survey data and ONS official statistics.
- Robustness/validation: included checks for knowledge-cutoff leakage and stability of validation answers over time.
Implications for AI Economics
- Useful tool for macro / survey research:
- LLMs can be used to generate synthetic surveys and to perform controlled treatment experiments (because the same synthetic agents can be re-run under alternate information states).
- Shapley-style decompositions provide interpretable attributions of prompt-content effects, enabling systematic sensitivity analysis of information treatments.
- Applications: stress-testing survey design, generating counterfactual public-opinion statistics, preliminary agent-based simulations when calibrated carefully.
- Calibration and constraints:
- Temperature and prompt engineering materially affect aggregate moments—practitioners should tune on out-of-sample/cross-validation data and report sensitivity.
- LLMs can reproduce population-level and demographic regularities, but their micro-level instability and internal inconsistencies limit their use for individual-level policy inferences or decision-making that requires coherent causal reasoning.
- Validation and OOS testing are essential:
- Rigorous out-of-sample setups (e.g., exploiting model training cut-offs) are necessary to detect memorization or leakage and to assess how models handle extremes beyond their training distribution.
- Results from LLM-based simulations must be validated against human data before being used for policy or research conclusions.
- Ethical and interpretability considerations:
- LLMs inherit and can amplify demographic biases; matching either human survey patterns or official statistics may involve trade-offs that need explicit handling depending on the research objective.
- Transparent reporting of prompts, temperatures, random seeds, and decomposition methods is required for reproducibility and for assessing fairness.
- Directions for future research:
- Test other LLM families and later versions (to assess how model updates and leakage change behavior).
- Enhance conditioning (richer information sets, dynamic histories) and investigate fine-tuning to instill more consistent economic world models.
- Extend interpretability tools to capture higher-order interactions and to better assess internal consistency across related economic questions.
- Evaluate how synthetic-agent results translate into real-world decision-making and policy design, especially where individual-level fidelity matters.
Short summary: GPT-3.5, prompted as synthetic IAS respondents and probed out-of-sample, can match aggregate and many demographic patterns in inflation attitudes and shows human-like overweighting of salient price components (notably food). Novel Shapley decompositions give interpretable attributions of prompt effects. But weak micro-level fidelity and internal inconsistencies mean LLMs are best used cautiously for population-level simulation, survey design, and exploratory analysis—always with empirical validation and transparency.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| GPT tracks aggregate survey projections and official statistics at short horizons. Decision Quality | positive | aggregate inflation projections (LLM outputs vs. survey projections and official statistics) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. Decision Quality | positive | household-level inflation perceptions by demographic subgroup |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. Other | positive | attribution of LLM output drivers to prompt content |
Reading fidelity
high
Study strength
medium
|
not reported
|
| GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. Decision Quality | positive | sensitivity of inflation perceptions to food inflation information |
Reading fidelity
high
Study strength
medium
|
not reported
|
| GPT lacks a consistent model of consumer price inflation. Decision Quality | negative | consistency/coherence of modelled consumer price inflation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The quasi-experimental design exploits GPT's training cut-off in September 2021, meaning it has no knowledge of the subsequent UK inflation surge. Other | null_result | model knowledge/availability of post-September 2021 inflation information |
Reading fidelity
high
Study strength
high
|
not reported
|
| More generally, this synthetic-survey approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design. Other | positive | utility of synthetic-survey approach for LLM evaluation and survey design |
Reading fidelity
high
Study strength
speculative
|
not reported
|