The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models mirror human biases on many economic preference tasks as they scale, but produce more rational belief judgments at higher capability levels; a simple prompt to 'make a rational decision' noticeably reduces bias across models.

Behavioral Economics of AI: LLM Biases and Corrections
Pietro Bini, Lin William Cong, Xing Huang, Lawrence J. Jin · February 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Pietro Bini unresolved corpus identity
  2. Lin William Cong unresolved corpus identity
  3. Xing Huang unresolved corpus identity
  4. Lawrence J. Jin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. P. Bini provider ID
  2. L. Cong provider ID
  3. Xin Huang provider ID
  4. Lawrence J. Jin provider ID
LLMs show systematic, task-dependent behavioral patterns—becoming more human-like on preference-based economic tasks as they scale, while advanced large models often give rational answers on belief-based tasks—and explicit prompting to 'be rational' substantially reduces exhibited biases.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.

Summary

Main Finding

Advanced and larger LLMs display systematic, task-dependent behavioral patterns: for preference-based tasks they become more human-like (and therefore more irrational by Expected Utility standards) as model size or recency increases, while for belief-based (statistical/forecasting) tasks they become more rational. A brief role‑priming instruction that asks the model to act as a rational Expected‑Utility investor reduces biases modestly; other attempted debiasing methods were largely ineffective.

Key Points

  • Five descriptive observations:
  • Preference tasks (prospect‑theory style choices, risk/time preferences): more advanced/larger models give answers closer to majority human responses and deviate from Expected Utility.
  • Belief tasks (forecasting, probability updating): more advanced/larger models give increasingly rational (Bayesian/statistically correct) responses.
  • Cross‑family heterogeneity: e.g., Gemini tends to be less rational/more human‑like on preferences vs GPT; Llama is less rational on beliefs vs GPT; Claude often similar to GPT.
  • Experimental‑economics tasks replicate these patterns: (a) in AR(1) forecasting (Afrouzi et al.), small advanced models overestimate persistence like humans, while large models estimate persistence nearer the truth; (b) in investment given price‑trajectories (Bose et al.), large models overweight visual salience (human‑like) more than smaller models.
  • Debiasing: role‑priming (ask model to be a rational Expected‑Utility investor before answering) reduces human‑like biases for both preference and belief tasks, via changes in reported confidence and reasoning mode; magnitude of correction is modest. Combining priming with extra factual information was ineffective.

  • Conjectured mechanisms:

    • Preference drift toward human patterns is plausibly driven by more RLHF in larger/advanced models (alignment to human feedback/preferences).
    • Improved rationality on belief tasks likely reflects larger training data and model capacity enabling better identification of statistical regularities.
  • Limitations noted by authors: results are descriptive, not causal; models and families evolve rapidly; some LLMs lacked graphical input so not all tasks were run on all models.

Data & Methods

  • Experimental stimuli:
    • Cognitive‑psychology benchmark questions (classic tasks documenting prospect‑theory preferences, overextrapolation, overconfidence). Each LLM response classified as: rational (Expected Utility/Bayesian), human‑like (majority human response that is irrational), or non‑human.
    • Experimental‑economics replications:
      • Afrouzi et al. (2023): AR(1) forecasting experiments (baseline, xt+1 & xt+5, and known‑AR variant). Measure perceived ρ implied by forecasts and compare to true ρ.
      • Bose et al. (2022): investment decisions after observing stock price trajectories; evaluate dependence on visual salience and other human‑identified factors.
  • Models tested (12 total across 4 families; cross‑sectional and time‑series comparisons):
    • OpenAI: GPT‑4 (benchmark), GPT‑4o (smaller variant), GPT‑3.5 Turbo (predecessor)
    • Anthropic: Claude 3 Opus (benchmark), Claude 3 Haiku (smaller), Claude 2 (predecessor)
    • Google: Gemini 1.5 Pro (benchmark), Gemini 1.5 Flash (smaller), Gemini 1.0 Pro (predecessor)
    • Meta: Llama 3 70B (benchmark), Llama 3 8B (smaller), Llama 2 70B (predecessor)
  • Data collection:
    • Responses collected via APIs; prompts adapted from original experiments to elicit comparable LLM answers. Some tasks required graphical inputs; six of twelve models did not support images so those tasks were not run on those models.
  • Analysis:
    • Compare LLM responses to rational benchmarks and majority human responses.
    • Infer perceived autoregressive coefficient (ˆρ) from sequences of forecasts.
    • Regress investment amounts on task features (e.g., salience) to test whether LLM decisions mirror human drivers.
    • Evaluate debiasing interventions (role priming; role priming + extra info) and measure changes in choice rationality, confidence, and self‑reported reasoning type.

Implications for AI Economics

  • For researchers using LLMs as social‑science tools:
    • Do not assume LLMs are neutral: they can systematically mirror human irrationalities in preference tasks while being more rational in belief tasks. Validate model behavior by task type before substituting LLMs for human subjects.
    • Distinguish belief vs preference evaluations when designing LLM‑based experiments or agent‑based models.
  • For financial and economic deployment of LLM agents:
    • Model choice and size matter: larger/advanced models can be more reliable in statistical forecasting but may reproduce human biases in preference‑sensitive decisions (investment choices, risk profiles).
    • Simple prompting (role priming toward rationality) can partially mitigate biases, but debiasing remains an open challenge—important for risk management, regulation, and auditability of AI‑assisted decisions.
  • For model development and governance:
    • Training objectives (e.g., RLHF aligning to human preferences) likely shape behavioral patterns; design choices should be explicit about intended normative behavior in economic contexts.
    • Systematic benchmarking (the authors propose a public database of experimental questions) is necessary for ongoing monitoring as models evolve.
  • Research directions:
    • Causal identification of mechanisms (RLHF vs data/scale) driving preference vs belief patterns.
    • More robust debiasing techniques beyond priming.
    • Broader, continuous evaluation across evolving LLM families, multimodal inputs, and real‑world financial tasks.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic, broad experimental evidence that LLMs exhibit consistent patterns of behavior on canonical behavioral-economics tasks across model families, versions, and scales, and shows robustness to prompting; however, findings are limited to model behavior in synthetic, text-prompted experiments rather than causal impacts on real-world economic outcomes, and are sensitive to prompt design, model updates, and deployment context. Methods Rigorhigh — Authors apply well-established cognitive-psychology and experimental-economics tasks, test multiple prominent LLM families and versions across scale, compare preference- and belief-based tasks, and include robustness checks (e.g., alternative prompts including explicit 'make a rational decision' instructions); the design mirrors human experimental protocols and isolates systematic patterns, though results still depend on prompt choices, hyperparameters, and the particular models sampled. SampleLarge-scale experimental evaluation of multiple prominent LLM families and versions (e.g., popular transformer-based models across generations and parameter scales) using a comprehensive suite of canonical behavioral-economics tasks originally designed for humans — including preference-based tasks (risk, time discounting, framing, ultimatum/choice experiments) and belief-based tasks (Bayesian updating, base-rate problems, probabilistic reasoning); each model was probed with multiple prompts and conditions (default vs. rationality-directed prompts) to assess systematic patterns across scale and version. (The paper emphasizes breadth across model families and scales rather than a single deployed system.) Themeshuman_ai_collab adoption GeneralizabilityResults are limited to the specific LLM families, versions, and parameter settings tested and may not hold for untested or future models., Experiments use text-only, artificial decision tasks rather than real-world, field economic behavior or deployed systems interacting with users., Model behavior is sensitive to prompt wording, temperature and decoding settings, and other inference-time hyperparameters which may vary in deployment., Training data cutoffs and proprietary pretraining regimes mean findings may not generalize across models with different training histories or multimodal inputs., Cultural and language limitations: tasks likely tested in one language and may not reflect cross-lingual behavior.

Claims (4)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions. Decision Quality positive presence and patterning of behavioral biases in economic/financial decision tasks
Reading fidelity high
Study strength medium
not reported
0.18
In preference-based tasks, LLM responses become more human-like as models become more advanced or larger. Decision Quality positive human-likeness of preference-based responses
Reading fidelity high
Study strength medium
not reported
0.18
In belief-based tasks, advanced large-scale models frequently generate rational responses. Decision Quality positive rationality of belief-based responses
Reading fidelity high
Study strength medium
not reported
0.18
Prompting LLMs to make rational decisions reduces behavioral biases. Error Rate positive level of bias (bias reduction) in LLM decisions
Reading fidelity high
Study strength medium
not reported
0.18

Notes