Simulated households driven by language models reveal that the wording, timing and source of tariff threats matter: immediate, high-rate, and coherent escalation messages raise average inflation and unemployment expectations, while uncertainty and unspecified rates increase disagreement. Central-bank explanations can tighten consensus even when they do not uniformly lower mean expectations, but these results are derived from agent-based simulations calibrated to survey data and should be treated as hypotheses for human testing.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.
Summary
Main Finding
A calibrated multi-agent system (MAS) that turns 300 Michigan Survey of Consumers households into persistent LLM-based agents reproduces key features of human survey responses to a real tariff announcement and shows that features of tariff-threat communication—timing, rate precision, semantic sequence, complexity, narrative, and sender identity—jointly shape both the mean and dispersion of household inflation and unemployment expectations. Central-bank messages delivered after a high-impact tariff threat can narrow disagreement (coordinate beliefs), but their effects on average expectations depend strongly on message content.
Key Points
-
Framework and validation
- Built a dynamic MAS of 300 Household Agents (personas and initial beliefs from the Michigan Survey of Consumers, December 2024) interacting with social-media-style posts over 18 simulated months; treatment delivered in month 7.
- Agents report point forecasts, subjective probability distributions, and open-ended explanations each month; agents also generate posts that can feed into others’ information sets via rolling memory.
- Calibration to an observed “Liberation Day” tariff announcement (benchmark T1: immediate 10% tariff) improves correspondence to human data on distributions, demographic patterns, and makes it harder for flexible classifiers to distinguish simulated from real responses.
- The authors emphasize a validation hierarchy: close benchmark correspondence permits “near-generalization” claims; more distant counterfactuals are scenario explorations and hypothesis generators, not causal identification for humans.
-
Experimental message design (12 arms)
- Control: placebo Trump post unrelated to macro variables.
- T1 (benchmark): immediate 10% universal tariff (based on Executive Order 14257).
- Variants T2–T12 manipulate implementation timing, announced rate/precision, semantic sequence (escalation vs reversals), message complexity (minimalist vs technical), narrative framing (MAGA vs tax-incidence), and sender identity (same text attributed to a Democratic leader).
- Trump-style posts were generated via an LLM role-playing a Trump Agent (few-shot examples) to preserve tone and expression.
-
Core empirical patterns from simulations
- Timing: Immediate implementation raises average point expectations more than delayed or uncertain timing; uncertain timing produces the greatest dispersion.
- Rate precision: Specified high rate (100%) raises average expectations more than 10% or unspecified rate; unspecified rate produces the most disagreement.
- Sequence coherence: Progressive escalation yields larger and more persistent revisions than a single high-rate message; frequent semantic reversals weaken responses.
- Message complexity: Minimalist wording produced larger average revisions than standard wording; technical/complex phrasing muted average revisions. Both unusually sparse and unusually complex messages increased dispersion versus a standard message.
- Narrative: MAGA-style and tax-incidence narratives attenuate average revisions (without materially reducing dispersion) but redirect causal attributions in agents’ open-ended explanations.
- Sender identity: The identical tariff wording attributed to a Democratic leader elicited a larger average response and greater dispersion than when attributed to the Republican leader (in these simulations).
- Mechanisms from open-ended answers: Effects were plausibly traced to attention, ambiguity, credibility/skepticism, political priors, and competing causal stories about who bears the tariff.
-
Central-bank communication experiment (post-escalation sequence)
- After the escalating tariff sequence (the most challenging environment), agents received one of five Fed-style messages: restating inflation objective; commitment to tighten if tariffs cause persistent broad-based price pressure; affirmation of independence; classification of tariff as a temporary relative-price shock; tolerance of a modest overshoot.
- Results: Messages framing persistent price pressure/tightening and messages tolerating overshoot produced the largest upward revisions in both inflation and unemployment expectations. Restating the objective and affirming independence had little average effect. The relative-price explanation was the only message that lowered point expectations. Crucially, all Fed messages reduced dispersion (improved coordination) even when they didn't lower the mean.
-
Limitations emphasized by authors
- MAS outputs are scenario-exploratory and hypothesis-generating; they do not substitute for human randomized experiments or provide causal estimates for people beyond the validated benchmark neighborhood.
- Results are sensitive to persona choice, pre-treatment priors, information-flow calibration, and LLM/model settings; each application requires fresh calibration and human validation.
Data & Methods
-
Data sources
- 300 human households sampled (stratified) from the Michigan Survey of Consumers (MSC); demographic variables and initial monthly expectations used to seed agent personas and priors.
- Social-media post texts on inflation and unemployment from X/Truth Social in December 2024 to shape pre-treatment information environment and as examples for style.
-
MAS construction and mechanics
- Each Household Agent is implemented with a large language model conditioned on the household’s observed characteristics and initial MSC-reported inflation/unemployment expectations.
- The environment is dynamic: agents see monthly information (treatment messages injected in month 7), maintain a rolling memory, interact via synthetic posts, and repeatedly answer surveys for 18 simulated months, producing a balanced panel.
- Elicitation: point forecasts, full subjective probability distributions, and open-ended rationale responses each period.
- Treatment arms: 1 control (placebo) + 12 tariff-threat variants (T1–T12). Treatments are delivered via LLM-generated presidential-style posts (Trump Agent) built by few-shot prompting with real posts as examples; T12 attributes identical wording to a Democratic leader.
- Calibration & validation: authors align the benchmark T1 MAS outputs to MSC responses collected April–November 2025. Validation checks include distributional comparisons, ability of flexible predictors to distinguish simulated vs human answers, and replication of demographic patterns (politics, income, gender).
-
Central-bank message experiment
- Conducted after agents experienced the progressively escalating tariff sequence (which produced the largest revisions).
- Five Fed messages chosen to represent distinct communication strategies (objective restatement; tightening commitment; independence statement; temporary relative-price framing; tolerance of overshoot).
- Outcomes: mean revisions and dispersion changes for inflation and unemployment expectations; open-ended explanations analyzed qualitatively.
-
Methodological notes
- Within-agent counterfactuals: same 300 agents used across arms so comparisons are within-agent and isolate information differences.
- Open-ended responses used to infer hypothesized cognitive channels (attention, credibility, ambiguity, causal models).
- Authors stress that calibrated MAS must be linked to human validation and that behavior resemblance is necessary but not sufficient for causal transport.
Implications for AI Economics
-
Tool for policy communication design
- MAS with persona-conditioned LLM agents provides a low-cost, rapid way to explore many message variants and anticipate qualitative directions of belief updating and dispersion changes before running costly human experiments.
- Can help policymakers distinguish between (a) messages that shift average expectations and (b) messages that improve coordination (reduce disagreement)—two different policy communication objectives.
-
Hypothesis generation and prioritization
- The system surfaces mechanism hypotheses (attention, ambiguity, sender credibility, causal framing) that can be prioritized for targeted human survey experiments and field trials, improving experimental design efficiency.
-
Modeling information environments
- Embedding persisting identities, rolling memory, and algorithmic-style exposure (social-media framing) permits studying dynamics of expectation formation that single-shot surveys and aggregate time-series models struggle to represent.
-
Cautions and governance
- Results are model- and calibration-dependent; MAS outputs must be validated against human data for the specific use case.
- Ethical considerations: such tools could be (mis)used to craft persuasive public messaging; researchers and policymakers should apply safeguards, transparency, and human-in-the-loop validation.
- Generalizability: framework is portable to other policy shocks (fiscal, energy, public health), but each application requires re-calibration of personas, priors, information flows, and validation.
-
Research agenda for AI economics
- Use MAS as an intermediate step between computational scenario exploration and costly human trials; iterate with human validation to refine LLM conditioning, memory models, and information-flow representations.
- Investigate how recommender-system exposure and endogenous message generation (agents posting) amplify or attenuate policy communication effects.
- Integrate MAS outputs with equilibrium and welfare analysis before issuing prescriptive policy advice.
Assessment
Claims (14)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The multi-agent system uses 300 household agents drawn from the Michigan Surveys of Consumers, with each agent assigned observed household characteristics and initial inflation and unemployment expectations. Other | null_result | Composition and initial beliefs of the simulated household sample |
Reading fidelity
high
Study strength
medium
|
n=300
|
| Calibrated household agents reproduce some distributional, political, income, and gender patterns observed in human survey data after the Liberation Day tariff announcement. Other | positive | Correspondence between simulated and human inflation and unemployment expectations |
Reading fidelity
high
Study strength
medium
|
n=300
|
| Immediate tariff implementation raises household point expectations more than implementation delayed by six months or an uncertain implementation date. Fiscal And Macroeconomic | positive | Point forecasts of inflation and unemployment |
Reading fidelity
high
Study strength
low
|
n=300
|
| Uncertain tariff implementation timing produces the greatest dispersion in household expectations among the timing treatments. Fiscal And Macroeconomic | positive | Dispersion of inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| A stated 100 percent tariff raises average expectations more than a stated 10 percent tariff or an unspecified tariff rate. Fiscal And Macroeconomic | positive | Average inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| An unspecified tariff rate produces the greatest disagreement in household expectations among the rate treatments. Fiscal And Macroeconomic | positive | Dispersion of inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| A progressively escalating sequence of tariff messages produces a larger and more persistent expectation response than a single high-rate message, while repeated semantic reversals weaken the response. Fiscal And Macroeconomic | mixed | Magnitude and persistence of revisions in inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| Message complexity affects average expectation revisions: minimalist wording produces a larger average revision than standard wording, while technical language produces a smaller revision; both unusually sparse and unusually complex messages widen dispersion relative to standard wording. Fiscal And Macroeconomic | mixed | Average revisions and dispersion of inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| MAGA and tax-incidence narratives attenuate the average expectation revision without materially reducing dispersion, while changing the causal explanations agents provide. Fiscal And Macroeconomic | mixed | Average expectation revisions, belief dispersion, and causal explanations |
Reading fidelity
high
Study strength
low
|
n=300
|
| The same tariff language attributed to the Democratic leader produces a larger average response and greater dispersion than when it is attributed to the Republican leader. Fiscal And Macroeconomic | positive | Average revisions and dispersion of inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| After a progressively escalating tariff sequence, central-bank messages that emphasize persistent price pressure and possible tightening, or explicitly tolerate an inflation overshoot, generate the largest upward revisions in inflation and unemployment expectations. Fiscal And Macroeconomic | positive | Inflation and unemployment expectations after central-bank communication |
Reading fidelity
high
Study strength
low
|
n=300
|
| Among the tested Federal Reserve messages, explaining the tariff as a temporary relative-price shock is the only strategy that lowers point expectations. Fiscal And Macroeconomic | negative | Point inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| Every tested Federal Reserve message reduces belief dispersion, even though the messages differ in their effects on average expectations. Fiscal And Macroeconomic | negative | Dispersion of inflation and unemployment expectations |
Reading fidelity
high
Study strength
low
|
n=300
|
| The model-generated open-ended explanations are mechanism hypotheses rather than independent evidence about human cognition. Ai Safety And Ethics | null_result | Validity and interpretability of simulated cognitive mechanisms |
Reading fidelity
high
Study strength
high
|
n=300
|