The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Simulated households driven by language models reveal that the wording, timing and source of tariff threats matter: immediate, high-rate, and coherent escalation messages raise average inflation and unemployment expectations, while uncertainty and unspecified rates increase disagreement. Central-bank explanations can tighten consensus even when they do not uniformly lower mean expectations, but these results are derived from agent-based simulations calibrated to survey data and should be treated as hypotheses for human testing.

Tariff Threats, Macroeconomic Expectations, and Policy Communication Strategies: Experiments Based on a Multi-Agent System
Jianhao Lin, Lexuan Sun, Yixin Yan · August 31, 2026
arxiv descriptive low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jianhao Lin unresolved corpus identity
  2. Lexuan Sun unresolved corpus identity
  3. Yixin Yan unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jianhao Lin provider ID
  2. Lexuan Sun provider ID
  3. Yixin Yan provider ID
Calibrated LLM-based household agents show that a tariff threat's timing, stated rate and certainty, message complexity, narrative framing, and sender identity jointly alter simulated inflation and unemployment expectations and their dispersion, while certain central-bank explanations reduce disagreement though only some change average expectations.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.

Summary

Main Finding

A calibrated multi-agent system (MAS) that turns 300 Michigan Survey of Consumers households into persistent LLM-based agents reproduces key features of human survey responses to a real tariff announcement and shows that features of tariff-threat communication—timing, rate precision, semantic sequence, complexity, narrative, and sender identity—jointly shape both the mean and dispersion of household inflation and unemployment expectations. Central-bank messages delivered after a high-impact tariff threat can narrow disagreement (coordinate beliefs), but their effects on average expectations depend strongly on message content.

Key Points

  • Framework and validation

    • Built a dynamic MAS of 300 Household Agents (personas and initial beliefs from the Michigan Survey of Consumers, December 2024) interacting with social-media-style posts over 18 simulated months; treatment delivered in month 7.
    • Agents report point forecasts, subjective probability distributions, and open-ended explanations each month; agents also generate posts that can feed into others’ information sets via rolling memory.
    • Calibration to an observed “Liberation Day” tariff announcement (benchmark T1: immediate 10% tariff) improves correspondence to human data on distributions, demographic patterns, and makes it harder for flexible classifiers to distinguish simulated from real responses.
    • The authors emphasize a validation hierarchy: close benchmark correspondence permits “near-generalization” claims; more distant counterfactuals are scenario explorations and hypothesis generators, not causal identification for humans.
  • Experimental message design (12 arms)

    • Control: placebo Trump post unrelated to macro variables.
    • T1 (benchmark): immediate 10% universal tariff (based on Executive Order 14257).
    • Variants T2–T12 manipulate implementation timing, announced rate/precision, semantic sequence (escalation vs reversals), message complexity (minimalist vs technical), narrative framing (MAGA vs tax-incidence), and sender identity (same text attributed to a Democratic leader).
    • Trump-style posts were generated via an LLM role-playing a Trump Agent (few-shot examples) to preserve tone and expression.
  • Core empirical patterns from simulations

    • Timing: Immediate implementation raises average point expectations more than delayed or uncertain timing; uncertain timing produces the greatest dispersion.
    • Rate precision: Specified high rate (100%) raises average expectations more than 10% or unspecified rate; unspecified rate produces the most disagreement.
    • Sequence coherence: Progressive escalation yields larger and more persistent revisions than a single high-rate message; frequent semantic reversals weaken responses.
    • Message complexity: Minimalist wording produced larger average revisions than standard wording; technical/complex phrasing muted average revisions. Both unusually sparse and unusually complex messages increased dispersion versus a standard message.
    • Narrative: MAGA-style and tax-incidence narratives attenuate average revisions (without materially reducing dispersion) but redirect causal attributions in agents’ open-ended explanations.
    • Sender identity: The identical tariff wording attributed to a Democratic leader elicited a larger average response and greater dispersion than when attributed to the Republican leader (in these simulations).
    • Mechanisms from open-ended answers: Effects were plausibly traced to attention, ambiguity, credibility/skepticism, political priors, and competing causal stories about who bears the tariff.
  • Central-bank communication experiment (post-escalation sequence)

    • After the escalating tariff sequence (the most challenging environment), agents received one of five Fed-style messages: restating inflation objective; commitment to tighten if tariffs cause persistent broad-based price pressure; affirmation of independence; classification of tariff as a temporary relative-price shock; tolerance of a modest overshoot.
    • Results: Messages framing persistent price pressure/tightening and messages tolerating overshoot produced the largest upward revisions in both inflation and unemployment expectations. Restating the objective and affirming independence had little average effect. The relative-price explanation was the only message that lowered point expectations. Crucially, all Fed messages reduced dispersion (improved coordination) even when they didn't lower the mean.
  • Limitations emphasized by authors

    • MAS outputs are scenario-exploratory and hypothesis-generating; they do not substitute for human randomized experiments or provide causal estimates for people beyond the validated benchmark neighborhood.
    • Results are sensitive to persona choice, pre-treatment priors, information-flow calibration, and LLM/model settings; each application requires fresh calibration and human validation.

Data & Methods

  • Data sources

    • 300 human households sampled (stratified) from the Michigan Survey of Consumers (MSC); demographic variables and initial monthly expectations used to seed agent personas and priors.
    • Social-media post texts on inflation and unemployment from X/Truth Social in December 2024 to shape pre-treatment information environment and as examples for style.
  • MAS construction and mechanics

    • Each Household Agent is implemented with a large language model conditioned on the household’s observed characteristics and initial MSC-reported inflation/unemployment expectations.
    • The environment is dynamic: agents see monthly information (treatment messages injected in month 7), maintain a rolling memory, interact via synthetic posts, and repeatedly answer surveys for 18 simulated months, producing a balanced panel.
    • Elicitation: point forecasts, full subjective probability distributions, and open-ended rationale responses each period.
    • Treatment arms: 1 control (placebo) + 12 tariff-threat variants (T1–T12). Treatments are delivered via LLM-generated presidential-style posts (Trump Agent) built by few-shot prompting with real posts as examples; T12 attributes identical wording to a Democratic leader.
    • Calibration & validation: authors align the benchmark T1 MAS outputs to MSC responses collected April–November 2025. Validation checks include distributional comparisons, ability of flexible predictors to distinguish simulated vs human answers, and replication of demographic patterns (politics, income, gender).
  • Central-bank message experiment

    • Conducted after agents experienced the progressively escalating tariff sequence (which produced the largest revisions).
    • Five Fed messages chosen to represent distinct communication strategies (objective restatement; tightening commitment; independence statement; temporary relative-price framing; tolerance of overshoot).
    • Outcomes: mean revisions and dispersion changes for inflation and unemployment expectations; open-ended explanations analyzed qualitatively.
  • Methodological notes

    • Within-agent counterfactuals: same 300 agents used across arms so comparisons are within-agent and isolate information differences.
    • Open-ended responses used to infer hypothesized cognitive channels (attention, credibility, ambiguity, causal models).
    • Authors stress that calibrated MAS must be linked to human validation and that behavior resemblance is necessary but not sufficient for causal transport.

Implications for AI Economics

  • Tool for policy communication design

    • MAS with persona-conditioned LLM agents provides a low-cost, rapid way to explore many message variants and anticipate qualitative directions of belief updating and dispersion changes before running costly human experiments.
    • Can help policymakers distinguish between (a) messages that shift average expectations and (b) messages that improve coordination (reduce disagreement)—two different policy communication objectives.
  • Hypothesis generation and prioritization

    • The system surfaces mechanism hypotheses (attention, ambiguity, sender credibility, causal framing) that can be prioritized for targeted human survey experiments and field trials, improving experimental design efficiency.
  • Modeling information environments

    • Embedding persisting identities, rolling memory, and algorithmic-style exposure (social-media framing) permits studying dynamics of expectation formation that single-shot surveys and aggregate time-series models struggle to represent.
  • Cautions and governance

    • Results are model- and calibration-dependent; MAS outputs must be validated against human data for the specific use case.
    • Ethical considerations: such tools could be (mis)used to craft persuasive public messaging; researchers and policymakers should apply safeguards, transparency, and human-in-the-loop validation.
    • Generalizability: framework is portable to other policy shocks (fiscal, energy, public health), but each application requires re-calibration of personas, priors, information flows, and validation.
  • Research agenda for AI economics

    • Use MAS as an intermediate step between computational scenario exploration and costly human trials; iterate with human validation to refine LLM conditioning, memory models, and information-flow representations.
    • Investigate how recommender-system exposure and endogenous message generation (agents posting) amplify or attenuate policy communication effects.
    • Integrate MAS outputs with equilibrium and welfare analysis before issuing prescriptive policy advice.

Assessment

Paper Typedescriptive Evidence Strengthlow — Findings come from simulations of LLM agents anchored to survey data rather than randomized interventions on real humans or structural identification; calibration and validation steps increase credibility for the simulated environment, but they do not establish causal effects in human populations. Methods Rigormedium — The design carefully constructs controlled language treatments, reuses identical calibrated personas for within-agent counterfactuals, and reports validation exercises against survey distributions and demographic patterns; however, results hinge on LLM modelling choices, prompt engineering, calibration scope, and untested assumptions about how simulated agents map to human cognition, limiting causal inference. Sample300 households drawn by stratified random sampling from the Michigan Survey of Consumers (December 2024), with observed demographics and reported inflation/unemployment expectations; social-media posts on X/Truth Social about inflation/unemployment used to construct information environment; 300 LLM Household Agents initialized with those personas and priors produce an 18-month simulated balanced panel with monthly elicited point forecasts, subjective distributions, and open-ended explanations. Themesgovernance human_ai_collab IdentificationWithin-agent counterfactual simulations using 300 LLM-based Household Agents calibrated to 2024 Michigan Survey of Consumers respondents and social-media texts; validation against observed MSC responses; causal claims about humans are treated as hypotheses from scenario exploration rather than identified human treatment effects. GeneralizabilitySimulated agents are not equivalent to real human subjects; behavioral resemblance does not establish external causal validity., Calibration targets a specific U.S. survey and a particular time period (post-2024 policy environment), limiting transfer across countries, cohorts, or policy contexts., Results depend on LLM architecture, prompt design, and memory/interaction rules; different models or settings may produce different outcomes., Information-flow model (recommendation algorithms, social-network dynamics) is stylized and may not capture real-world exposure heterogeneity., Small computational sample (300 agents) and re-use of identical agents across arms may understate variability present in human experiments., Findings focus on tariff threats and central-bank messaging; not a substitute for equilibrium or welfare analysis of policy consequences.

Claims (14)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The multi-agent system uses 300 household agents drawn from the Michigan Surveys of Consumers, with each agent assigned observed household characteristics and initial inflation and unemployment expectations. Other null_result Composition and initial beliefs of the simulated household sample
Reading fidelity high
Study strength medium
n=300
0.18
Calibrated household agents reproduce some distributional, political, income, and gender patterns observed in human survey data after the Liberation Day tariff announcement. Other positive Correspondence between simulated and human inflation and unemployment expectations
Reading fidelity high
Study strength medium
n=300
0.18
Immediate tariff implementation raises household point expectations more than implementation delayed by six months or an uncertain implementation date. Fiscal And Macroeconomic positive Point forecasts of inflation and unemployment
Reading fidelity high
Study strength low
n=300
0.09
Uncertain tariff implementation timing produces the greatest dispersion in household expectations among the timing treatments. Fiscal And Macroeconomic positive Dispersion of inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
A stated 100 percent tariff raises average expectations more than a stated 10 percent tariff or an unspecified tariff rate. Fiscal And Macroeconomic positive Average inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
An unspecified tariff rate produces the greatest disagreement in household expectations among the rate treatments. Fiscal And Macroeconomic positive Dispersion of inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
A progressively escalating sequence of tariff messages produces a larger and more persistent expectation response than a single high-rate message, while repeated semantic reversals weaken the response. Fiscal And Macroeconomic mixed Magnitude and persistence of revisions in inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
Message complexity affects average expectation revisions: minimalist wording produces a larger average revision than standard wording, while technical language produces a smaller revision; both unusually sparse and unusually complex messages widen dispersion relative to standard wording. Fiscal And Macroeconomic mixed Average revisions and dispersion of inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
MAGA and tax-incidence narratives attenuate the average expectation revision without materially reducing dispersion, while changing the causal explanations agents provide. Fiscal And Macroeconomic mixed Average expectation revisions, belief dispersion, and causal explanations
Reading fidelity high
Study strength low
n=300
0.09
The same tariff language attributed to the Democratic leader produces a larger average response and greater dispersion than when it is attributed to the Republican leader. Fiscal And Macroeconomic positive Average revisions and dispersion of inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
After a progressively escalating tariff sequence, central-bank messages that emphasize persistent price pressure and possible tightening, or explicitly tolerate an inflation overshoot, generate the largest upward revisions in inflation and unemployment expectations. Fiscal And Macroeconomic positive Inflation and unemployment expectations after central-bank communication
Reading fidelity high
Study strength low
n=300
0.09
Among the tested Federal Reserve messages, explaining the tariff as a temporary relative-price shock is the only strategy that lowers point expectations. Fiscal And Macroeconomic negative Point inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
Every tested Federal Reserve message reduces belief dispersion, even though the messages differ in their effects on average expectations. Fiscal And Macroeconomic negative Dispersion of inflation and unemployment expectations
Reading fidelity high
Study strength low
n=300
0.09
The model-generated open-ended explanations are mechanism hypotheses rather than independent evidence about human cognition. Ai Safety And Ethics null_result Validity and interpretability of simulated cognitive mechanisms
Reading fidelity high
Study strength high
n=300
0.3

Notes