4 cumulative citations
View corpus contextLLM traders mimic some—but not all—behavioral-finance patterns in simulated markets; prompted agents switch strategies in theoretically expected directions only intermittently, implying current LLM-based simulations are informative but not yet reliable proxies for real traders.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. Investors are typically classified as fundamental or technical traders, but most simulations fix strategies at initialization, failing to reflect real-world trading dynamics. In this work, we assess whether agents' strategy switching aligns with financial theory, providing a framework for this evaluation. We operationalize four behavioral-finance drivers-loss aversion, herding, wealth differentiation, and price misalignment-as personality traits set via prompting and stored long-term. In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Our results show that recent LLMs' switching behavior is only partially consistent with behavioral-finance theories, highlighting the need for further refinement in aligning agent behavior with financial theory.
Summary
Main Finding
LLM-based agents in a year-long stock-market simulation display only partial behavioral consistency with standard behavioral-finance theories. Across four tested LLMs (GPT-4o-mini, Gemini-2.5, Deepseek-Chat, Qwen-2.5), agents reliably exhibit loss-aversion-like persistence in staying with a current trading style after losses, but alignment is weaker or inconsistent for herding, wealth-differentiation, and price-misalignment drivers. The paper develops an operational framework (persona prompts + long-term memory + counterfactual ledgers) and formal alignment metrics to evaluate whether LLM agents’ style-switching maps to theory.
Key Points
- Problem: Prior LLM-agent market simulations typically fix agent styles (fundamental vs. technical). Real traders switch styles; the paper tests whether LLM agents switch in ways consistent with behavioral finance.
- Behavioral drivers encoded as persistent persona traits (via prompts): loss aversion, herding tendency, wealth-differentiation sensitivity, and price-misalignment sensitivity.
- Agent population: factorial design (2^5 = 32 agents) crossing four binary trait flags and an initial style (Tech vs Fund), so aligned vs. non-aligned cohorts (n=16 each per factor).
- Simulation: full 2024 trading year (253 days) on 5 representative S&P 500 stocks (MSFT, ICE, VRTX, CAT, CLX); daily trading decisions; style review every 10 trading days; each agent records an actual ledger and a counterfactual ledger for the alternate style.
- Evaluation: four alignment metrics (Loss-Aversion Alignment Score (LAS), Herd-Alignment Score (HAS), Advantage/Wealth-Diff alignment, Mispricing alignment); comparisons between aligned and non-aligned cohorts via one-sided Mann–Whitney U tests with effect sizes (rank-biserial, Cliff’s δ, CLES).
- Main empirical result: strong and consistent alignment with loss-aversion across models (significant U-tests and moderate-to-large effect sizes). Herding, wealth-differentiation, and mispricing alignment show mixed or weak evidence depending on the model; only some model/factor pairs reach significance.
- Example stats (selected): GPT-4o-mini Loss Aversion U=214, p=0.0003 (large effect); GPT-4o-mini Wealth Differentiation U=174, p=0.04 (weaker effect); Herding and Price Misalignment nonsignificant for most models (exceptions: Qwen shows some herding signal).
- Explainability: style-switch decisions are accompanied by natural-language rationales from agents, enabling traceability of switching decisions.
Data & Methods
- Data
- Price–volume: daily split-adjusted close, prior close, change, pct_chg, volume, N-day moving averages (N=5,10,30) for each selected ticker across 2024.
- Fundamentals: quarterly disclosures used to derive Leverage (Debt Ratio), Current Ratio, Operating Cash Flow (OCF), Free Cash Flow (FCF).
- Stock pool: one representative per five sectors (MSFT, ICE, VRTX, CAT, CLX) to eliminate random cross-sectional variation and isolate behavior-driven outcomes.
- Agent design
- 32 agents via full-factorial design over five binary factors: (loss aversion present/neutral), (herding present/neutral), (wealth-diff sens present/neutral), (mispricing sens present/neutral), (initial style Tech/Fund).
- Initial wealth equal for all, 50% allocated to equities (10% to each stock), long-only, no leverage.
- Memory module: persistent persona traits, actual and counterfactual ledgers, daily updates plus block-level summaries every 10 trading days.
- Trading mechanics
- Daily decision: agent receives up-to-date price/volume/indicator data and chooses Buy/Sell/Hold for each stock (actual and simulated counterfactual under opposite style).
- Execution: orders executed at market open using split-adjusted prices; ledgers updated accordingly.
- Style review: every 10 trading days agents evaluate block P&L, counterfactual performance, population style shares, and persona traits to decide whether to switch style (decision rationale produced in natural language).
- Metrics & statistical tests
- LAS (aggregate stay-after-loss magnitude), HAS (fraction of others using alternative style at switches), Advantage score (wealth-diff alignment), Mispricing score (movement toward fundamental style when price deviates from estimated value).
- Compare aligned vs non-aligned cohorts using one-sided Mann–Whitney U tests (aligned > non-aligned). Report U, one-sided p-value, and nonparametric effect sizes (CLES, rank-biserial r_rb, Cliff’s δ).
- Models tested
- GPT-4o-mini, Gemini-2.5-flash-lite-thinking-8192, Deepseek-Chat, Qwen-2.5-72B-Instruct.
Implications for AI Economics
- Feasibility of LLM Agents for Behavioral ABM: The study demonstrates a practical pipeline to embed behavioral traits into LLM agents (prompts + memory + counterfactual evaluation) and evaluate micro-level consistency with economic theory. This advances the methodological toolkit for AI-driven agent-based models in economics.
- Partial validity cautions for macro inference: Because LLM agents align strongly with loss aversion but only partially with other behavioral drivers, macro phenomena produced by LLM-populated markets (e.g., volatility, bubbles) may be biased if simulations assume fully theory-consistent switching behavior. Modelers should validate micro-level behaviors before extrapolating to system-level claims.
- Need for model and prompt calibration: Differences across backbone LLMs (and across drivers) indicate that behavioral fidelity depends on model choice and prompt engineering. For economic ABMs intended to reflect human-like switching, careful calibration (and reporting) of persona prompts and validation metrics is necessary.
- Methodological contribution: The alignment metrics and counterfactual-ledger approach provide a replicable evaluation framework to test other behavioral hypotheses (risk preferences, overconfidence, news sensitivity) in LLM-agent simulations.
- Directions for future work to strengthen behavioral realism:
- Broader market realism: larger, more diverse stock universes, order-book microstructure, latency/noise, and transaction costs.
- richer social dynamics: explicit communication networks, private vs public signals, and information frictions to more fully realize herding channels.
- Data-driven calibration: match LLM agent switching rates to human-subject experiments or historical trader behavior to tune prompt priors and switching thresholds.
- Robustness and generalization: test additional LLM architectures, longer horizons, stochastic shocks, and finer decision cadences.
- Practical implication for policy and research: Results suggest LLM agents can be used to test some behavioral mechanisms, but regulators and researchers should treat positive simulation outcomes (e.g., emergence of a bubble) as hypothesis-generating rather than conclusive without micro-level behavioral validation.
If you want, I can: - Extract and report the full Table 3 results in a compact table or CSV, - Produce example persona prompts used for each trait (from the paper’s Appendix), - Propose an experimental plan to improve alignment for the weaker drivers (herding, mispricing).
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. Adoption Rate | positive | use of LLM agents in financial market simulations (adoption) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Most simulations fix trading strategies at initialization, failing to reflect real-world trading dynamics. Other | negative | design choice in prior simulations (strategy fixation) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We operationalize four behavioral-finance drivers—loss aversion, herding, wealth differentiation, and price misalignment—as personality traits set via prompting and stored long-term. Other | positive | encoding of behavioral drivers into agent prompts (trait operationalization) |
Reading fidelity
high
Study strength
high
|
not reported
|
| In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. Other | positive | agent behavior over time (strategy reassessment frequency) |
Reading fidelity
high
Study strength
high
|
not reported
|
| We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Decision Quality | positive | alignment between agent style-switching and financial-theory-predicted behavior (measured by the introduced metrics) |
Reading fidelity
high
Study strength
high
|
not reported
|
| Recent LLMs' switching behavior is only partially consistent with behavioral-finance theories. Decision Quality | mixed | consistency/alignment of agents' strategy-switching behavior with behavioral-finance theories |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The partial consistency of LLM agents' switching behavior highlights the need for further refinement in aligning agent behavior with financial theory. Other | positive | requirement for methodological refinement to improve alignment |
Reading fidelity
high
Study strength
medium
|
not reported
|