The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

GPT-3.5 mirrors UK household inflation perceptions at short horizons and across demographic groups, yet fails to internalize the post-2021 inflation surge because of its training cutoff and shows no consistent model of consumer-price dynamics.

Inflation Attitudes of Large Language Models
Nikoleta Anesti, Edward Hill, Andreas Joseph · December 16, 2025
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Nikoleta Anesti unresolved corpus identity
  2. Edward Hill unresolved corpus identity
  3. Andreas Joseph unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Nikoleta Anesti provider ID
  2. Edward Hill provider ID
  3. A. Joseph provider ID
GPT-3.5-turbo, when prompted to mimic a household inflation survey, tracks aggregate short-horizon inflation perceptions and reproduces key demographic patterns (income, tenure, class) but lacks a consistent internal model of consumer-price dynamics and misses post-2021 inflation due to its training cutoff.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.

Summary

Main Finding

GPT-3.5-turbo (gpt-3.5-turbo-0613), when prompted to act as households in a synthetic version of the Bank of England’s Inflation Attitudes Survey (IAS), can reproduce aggregate survey distributions and many demographic regularities in inflation perceptions and short-horizon expectations. A Shapley-value decomposition adapted to the synthetic-survey setting shows prompt-conditioned drivers of responses (notably a human-like oversensitivity to food inflation). However, the model lacks a coherent internal model of consumer-price inflation: micro-level correspondence with individual human responses is weak and partially unstable, and the LLM exhibits internal inconsistencies that limit its use for individual-level inference.

Key Points

  • Quasi-experimental design: exploited GPT’s training cut-off (Sept 2021) to probe out-of-sample behaviour during the UK inflation surge (peaking late 2022).
  • Synthetic surveys: GPT was prompted with synthetic personas (demographics) and economic information (subcomponent inflation) to answer IAS questions on perceived current inflation and expectations at 1, 2 and 5 years.
  • Aggregate fit: with temperature tuning, GPT can match aggregate response distributions and short-horizon expectations fairly well.
  • Meso-level alignment: GPT replicates key demographic correlations observed in real IAS responses (income, housing tenure, social class) and often aligns more closely with official statistics than with human survey responses.
  • Component sensitivity: GPT shows heightened sensitivity to salient components (food/restaurant inflation), similar to documented human biases.
  • Micro-level limits: individual-level responses are unstable across runs and inconsistent in ways that indicate GPT does not have a consistent world model of CPI dynamics.
  • Explainability: introduced a Shapley-value decomposition for LLM outputs in the synthetic-survey setting to quantify contributions of prompt bits (demographics, each inflation component) and their interactions.
  • Practical cautions: results depend on prompt design, temperature, and the specific model/version; ethical concerns and demographic biases persist. The model is used as a representative testbed, not an endorsement of any vendor.

Data & Methods

  • Model and API: OpenAI chat completions API using gpt-3.5-turbo-0613. System + user prompts specified a persona (age, gender, region, income, education, tenure, social class) and economic conditioning.
  • Economic conditioning: three-month averages of year-on-year inflation for CPIH subcomponents immediately preceding the survey date were given in prompts:
    • Cross-validation sample (Nov 2022): food 15.0%, restaurants 8.1%, energy 76.0%, other 6.0%, CPIH 9.3%.
    • Main test sample (Feb 2023): food 17.0%, restaurants 9.8%, energy 88.0%, other 5.0%, CPIH 9.2%.
    • These values are well outside GPT’s training-range (pre-Oct 2021 averages: CPIH ≈ 2.6%).
  • Experimental design:
    • Two samples: 2022Q4 (Nov 2022) used for cross-validation/tuning; 2023Q1 (Feb 2023) as main test.
    • Temperature grid: T ∈ {0, 0.25, 0.5, 0.75, 1, 1.25, 1.5} to control deterministic vs. stochastic outputs; temperature tuned on CV sample to match aggregate moments.
    • Randomization: response option order was scrambled per respondent to avoid order bias.
    • Treatments: presence vs omission (or baseline historic averages) of economic-condition bits in prompts defined treated vs control states; because the same synthetic subject can be re-run under different treatments, individual treatment effects were computed directly.
  • Explainability: applied a Shapley-value decomposition (adapted for discrete prompt-treatment bits) to attribute contributions of each prompt element (e.g., food inflation, energy inflation, demographics) and their interactions to the model’s chosen survey response.
  • Benchmarks: compared synthetic GPT responses at micro (individual), meso (demographic groups), and macro (aggregate distributions) levels against actual IAS survey data and ONS official statistics.
  • Robustness/validation: included checks for knowledge-cutoff leakage and stability of validation answers over time.

Implications for AI Economics

  • Useful tool for macro / survey research:
    • LLMs can be used to generate synthetic surveys and to perform controlled treatment experiments (because the same synthetic agents can be re-run under alternate information states).
    • Shapley-style decompositions provide interpretable attributions of prompt-content effects, enabling systematic sensitivity analysis of information treatments.
    • Applications: stress-testing survey design, generating counterfactual public-opinion statistics, preliminary agent-based simulations when calibrated carefully.
  • Calibration and constraints:
    • Temperature and prompt engineering materially affect aggregate moments—practitioners should tune on out-of-sample/cross-validation data and report sensitivity.
    • LLMs can reproduce population-level and demographic regularities, but their micro-level instability and internal inconsistencies limit their use for individual-level policy inferences or decision-making that requires coherent causal reasoning.
  • Validation and OOS testing are essential:
    • Rigorous out-of-sample setups (e.g., exploiting model training cut-offs) are necessary to detect memorization or leakage and to assess how models handle extremes beyond their training distribution.
    • Results from LLM-based simulations must be validated against human data before being used for policy or research conclusions.
  • Ethical and interpretability considerations:
    • LLMs inherit and can amplify demographic biases; matching either human survey patterns or official statistics may involve trade-offs that need explicit handling depending on the research objective.
    • Transparent reporting of prompts, temperatures, random seeds, and decomposition methods is required for reproducibility and for assessing fairness.
  • Directions for future research:
    • Test other LLM families and later versions (to assess how model updates and leakage change behavior).
    • Enhance conditioning (richer information sets, dynamic histories) and investigate fine-tuning to instill more consistent economic world models.
    • Extend interpretability tools to capture higher-order interactions and to better assess internal consistency across related economic questions.
    • Evaluate how synthetic-agent results translate into real-world decision-making and policy design, especially where individual-level fidelity matters.

Short summary: GPT-3.5, prompted as synthetic IAS respondents and probed out-of-sample, can match aggregate and many demographic patterns in inflation attitudes and shows human-like overweighting of salient price components (notably food). Novel Shapley decompositions give interpretable attributions of prompt effects. But weak micro-level fidelity and internal inconsistencies mean LLMs are best used cautiously for population-level simulation, survey design, and exploratory analysis—always with empirical validation and transparency.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper provides strong descriptive and comparative evidence that GPT-3.5 reproduces several aggregate and demographic patterns in inflation perceptions and is sensitive to specific price signals, leveraging a clear quasi-experimental timing advantage; however, it does not establish causal effects of LLMs on economic outcomes, results hinge on prompt design and a single model/version, and the training-cutoff strategy tests information absence rather than causal mechanisms about economic behavior. Methods Rigorhigh — Careful matching to an official survey (IAS), explicit exploitation of a clear exogenous training cutoff, pre/post comparisons to official statistics, and the novel application of Shapley-value decomposition to attribute prompt-driven influences together constitute a rigorous and transparent analytical approach; remaining methodological caveats (prompt sensitivity, single-model focus, and external validity) are acknowledged and tested where feasible. SampleSynthetic survey responses generated from OpenAI GPT-3.5-turbo prompted to mimic the Bank of England Inflation Attitudes Survey across demographic cells (income, housing tenure, social class, etc.), compared to actual IAS household survey data and official UK inflation statistics (pre- and post-September 2021); analysis leverages temporal variation around GPT's training cutoff and disaggregated demographic cells rather than individual-level longitudinal respondent panels. Themeshuman_ai_collab adoption IdentificationExploit GPT-3.5-turbo's fixed training cutoff (September 2021) as a natural experiment: compare model-generated inflation perceptions (using prompts that mimic the Bank of England's Inflation Attitudes Survey and matched demographic attributes) to contemporaneous household survey responses and official statistics before and after the post-2021 UK inflation surge; use timing variation to show which signals the model has encoded and a Shapley-value prompt decomposition to identify drivers of model outputs. GeneralizabilityFindings are specific to GPT-3.5-turbo and the particular model snapshot (training cutoff Sept 2021); newer or different LLMs may behave differently., Results depend on prompt design and the synthetic-survey protocol; alternative prompts or sampling schemes could change outcomes., Geographic focus on UK inflation limits transferability to other countries or price regimes., The training-cutoff strategy tests information absence rather than how models update or how humans interact with models in real-world settings., Does not measure downstream economic impacts (e.g., on behavior, markets, or policy), so limited for causal inference on macro/economic outcomes.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
GPT tracks aggregate survey projections and official statistics at short horizons. Decision Quality positive aggregate inflation projections (LLM outputs vs. survey projections and official statistics)
Reading fidelity high
Study strength medium
not reported
0.48
At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. Decision Quality positive household-level inflation perceptions by demographic subgroup
Reading fidelity high
Study strength medium
not reported
0.48
A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. Other positive attribution of LLM output drivers to prompt content
Reading fidelity high
Study strength medium
not reported
0.48
GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. Decision Quality positive sensitivity of inflation perceptions to food inflation information
Reading fidelity high
Study strength medium
not reported
0.48
GPT lacks a consistent model of consumer price inflation. Decision Quality negative consistency/coherence of modelled consumer price inflation
Reading fidelity high
Study strength medium
not reported
0.48
The quasi-experimental design exploits GPT's training cut-off in September 2021, meaning it has no knowledge of the subsequent UK inflation surge. Other null_result model knowledge/availability of post-September 2021 inflation information
Reading fidelity high
Study strength high
not reported
0.8
More generally, this synthetic-survey approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design. Other positive utility of synthetic-survey approach for LLM evaluation and survey design
Reading fidelity high
Study strength speculative
not reported
0.08

Notes