2 cumulative citations
View corpus contextDeliberation prompts make large language models more rational but also more pliable: crude in‑context priming pushes models to extreme emotion-driven choices, while representation-level steering produces subtler, less reliable human-like biases, highlighting a trade-off between controllability and human-aligned behavior.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large Language Models (LLMs) are increasingly positioned as decision engines for hiring, healthcare, and economic judgment, yet real-world human judgment reflects a balance between rational deliberation and emotion-driven bias. If LLMs are to participate in high-stakes decisions or serve as models of human behavior, it is critical to assess whether they exhibit analogous patterns of (ir)rationalities and biases. To this end, we evaluate multiple LLM families on (i) benchmarks testing core axioms of rational choice and (ii) classic decision domains from behavioral economics and social norms where emotions are known to shape judgment and choice. Across settings, we show that deliberate "thinking" reliably improves rationality and pushes models toward expected-value maximization. To probe human-like affective distortions and their interaction with reasoning, we use two emotion-steering methods: in-context priming (ICP) and representation-level steering (RLS). ICP induces strong directional shifts that are often extreme and difficult to calibrate, whereas RLS produces more psychologically plausible patterns but with lower reliability. Our results suggest that the same mechanisms that improve rationality also amplify sensitivity to affective interventions, and that different steering methods trade off controllability against human-aligned behavior. Overall, this points to a tension between reasoning and affective steering, with implications for both human simulation and the safe deployment of LLM-based decision systems.
Summary
Main Finding
Reasoning-enabled LLMs (models that generate intermediate “thinking” tokens) are substantially more internally consistent with classical rational-choice axioms and tend to default to expected-value (EV) maximization. However, the same reasoning mechanisms also amplify susceptibility to affective steering. Two emotion-steering methods produce qualitatively different effects: in‑context priming (ICP) yields strong, often extreme directional shifts (hard to calibrate), while representation‑level steering (RLS) produces more psychologically plausible, dose‑responsive but weaker and less reliable changes. This creates a tension between using LLMs to (a) simulate human-like, emotion-driven biases and (b) deploy LLMs as stable, unbiased decision agents.
Key Points
- Reasoning vs non-reasoning:
- “Thinking” LLMs show higher compliance with rational‑choice axioms (completeness, transitivity, continuity, independence) than non‑thinking models.
- Thinking models, absent explicit instruction otherwise, default toward EV‑maximizing choices and near-linear probability weighting.
- Emotion steering methods:
- ICP (lexical/persona priming in prompt) produces large, often extreme, directional shifts aligned with human emotion effects (e.g., fear → strong risk aversion; anger → more risk taking) but is hard to control and can collapse model behavior (e.g., fear ICP nearly eliminates gambling).
- RLS (low‑rank vector injections into hidden activations) yields smaller, smoother, and more dose‑dependent effects that better match psychological theory (e.g., graded increases in risk aversion under fear) but with lower reliability across instances.
- Interaction of reasoning and affect:
- The mechanisms that increase rationality also let the model “rationalize” induced affect—reasoning traces often incorporate the steered emotional state, magnifying behavioral changes.
- Emotional steering modestly reduces axiom-consistency scores; stronger steering → larger drops in rationality compliance.
- Domain‑specific deviations from human benchmarks:
- Weak or absent human-like loss aversion in baseline thinking models, but ICP can induce large loss‑aversion-like geometry (sometimes exaggerated; e.g., reported fear‑ICP λ ≈ 2.94).
- Near‑complete ambiguity aversion in neutral thinking models; sadness‑RLS reduced ambiguity aversion (human‑aligned), whereas sadness‑ICP unexpectedly increased it.
- Reversed endowment effect (models sometimes price “owned” items lower than “not‑owned” ones), limited monetary depreciation representation, and other departures from canonical human patterns.
- Practical safety/robustness concerns:
- Reasoning modes improve normative behavior but increase vulnerability to emotional prompt injection and steering; trade‑off between fidelity to human behavior and safe, stable decision-making.
Data & Methods
- Models evaluated:
- Multiple LLM families and scales, including Llama-3.x variants, Olmo-3-7B, and Qwen-3 (notably Qwen-3-4B and Qwen-3-8B). Analyses focus on reasoning-enabled models (referred to as “thinking” models).
- Steering interventions:
- In‑Context Priming (ICP): emotion personas, vignettes, and lexical descriptors inserted into prompts; steering strength controlled by prompt wording.
- Representation‑Level Steering (RLS): low‑rank vector(s) injected into hidden activations; steering strength controlled by injection coefficient (e.g., scales tested: 20, 30, 35, 40).
- Rationality benchmark:
- Diagnostic battery testing four axioms required for Expected Utility representation: completeness, transitivity, continuity, independence.
- Scores summarize the proportion of decision instances satisfying axioms.
- Decision tasks / domains:
- Choices under risk (sure vs lottery), probability‑weighting estimation (Prelec parametric form w(p; α, β)), utility curvature over gains (u(x) = x^ρ), ambiguity aversion (known vs unknown probabilities), loss aversion (λ estimation), temporal discounting, endowment effect, social norms, moral judgment, fairness/prosocial choices.
- Parametric estimation:
- Risk choice modeled with logistic link: Pr(risky) = σ(τ ΔU + b).
- For probability weighting: Prelec function fitted to certainty equivalents to estimate α and β.
- Utility curvature ρ and loss‑aversion λ estimated from fitted choice boundaries/iso‑utility lines.
- Key quantitative observations reported (examples from figures):
- Neutral (thinking) Qwen3: near-linear probability weighting (α ≈ 1.00, β ≈ 1.03) and near-linear utility (ρ ≈ 1.00).
- Fear via ICP: extreme distortions (example Prelec fit: α ≈ 0.66, β ≈ 6.00; utility ρ ≈ 0.20) — strong underweighting of objective probabilities and high curvature.
- Fear via RLS: milder effect (e.g., α ≈ 1.14, β ≈ 1.30; ρ ≈ 0.95) — modest increase in risk aversion with dose‑response behavior.
- Experimental controls:
- Steering strength varied; comparisons made across steering methods, strengths, and with neutral baselines. Appendix in original work reports additional robustness checks and results across prompt templates and other LLMs.
Implications for AI Economics
- Using LLMs as decision agents or prescriptive tools:
- Reasoning-enabled LLMs produce more coherent, EV‑oriented behavior and are preferable for tasks demanding internal consistency. But they are more susceptible to manipulation via emotional steering, raising deployment safety concerns in high‑stakes domains (hiring, healthcare, policy advice).
- Systems that must avoid unintended bias should either (a) remove or tightly control “thinking” chains when exposed to user text that could impart affective cues, or (b) include robust adversarial detection/mitigation for emotional prompt injection.
- Modeling human behavior and synthetic agents for experiments:
- If the goal is faithful simulation of human emotional biases, RLS is a better tool: it yields more psychologically plausible, graded effects that are easier to interpret and calibrate statistically.
- ICP can produce directionally correct but exaggerated human‑like responses; using ICP to generate synthetic human data risks producing nonrepresentative distributions (over‑extreme risk aversion, loss aversion, etc.). Researchers should validate and calibrate ICP outputs against human anchors.
- Policy simulation and economic forecasting:
- LLMs defaulting to EV maximization under thinking mode implies that model‑based simulations will understate human‑behavioral frictions (e.g., risk, loss aversion, ambiguity effects) unless explicitly and carefully steered. Policymakers using LLMs to predict behavioral responses should incorporate calibrated steers (preferably RLS) or hybrid models combining LLMs with behavioral parameterizations.
- Experimental design & best practices:
- Always report reasoning mode and steering treatment when using LLMs as agents in economic experiments.
- Validate steering effects across multiple strengths and steering mechanisms; prefer representation-level steering for reproducible, dose‑dependent behavior when trying to mimic human psychology.
- Include rationality/consistency diagnostics (axiom battery) as routine checks before using model outputs for downstream economic conclusions.
- Safety and governance:
- There is a trade‑off between aligning models to human biases (useful for simulation/empirical social‑science research) and ensuring robustness/unbiased decision-making in deployed systems. Governance frameworks should recognize this dichotomy and require distinct validation regimes depending on the intended use (simulation vs operational decision support).
- Future research directions for AI economics:
- Develop calibrated, reliable methods to induce human-like affective states in LLMs without causing catastrophic behavior collapse (ICP calibration, improved RLS reliability).
- Integrate explicit behavioral parameter modules (loss aversion, probability weighting) with LLM reasoning to combine interpretability and robustness.
- Study multi-agent interactions where some agents are reasoning LLMs and others are human or non‑thinking models to map how amplified susceptibility to steering propagates in markets or policy contexts.
Short actionable recommendations - For simulation of human behavior: prefer RLS with systematic calibration to human data; verify dose‑response and reliability. - For decision-support deployments: enable reasoning for consistency but add guardrails against emotional steering (monitoring, prompt sanitization, adversarial testing). - For researchers: always run the rationality axiom battery and report steering method and strength when using LLMs as behavioral agents.
If you want, I can extract the paper’s key quantitative parameter estimates (Prelec α/β, utility ρ, loss‑aversion λ) into a compact table, or draft a checklist for safe use of reasoning LLMs in economic decision systems.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Large Language Models (LLMs) are increasingly positioned as decision engines for hiring, healthcare, and economic judgment. Adoption Rate | positive | deployment/adoption of LLMs in decision-making domains |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We evaluate multiple LLM families on (i) benchmarks testing core axioms of rational choice and (ii) classic decision domains from behavioral economics and social norms where emotions shape judgment and choice. Decision Quality | null_result | performance on rational-choice and behavioral economics benchmarks |
Reading fidelity
high
Study strength
high
|
not reported
|
| Deliberate 'thinking' reliably improves rationality and pushes models toward expected-value maximization. Decision Quality | positive | rationality (adherence to rational-choice axioms) / expected-value maximization |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In-context priming (ICP) induces strong directional shifts in model choices that are often extreme and difficult to calibrate. Decision Quality | mixed | directional change in choices/responses due to emotion priming |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Representation-level steering (RLS) produces more psychologically plausible affective patterns in LLM decisions but with lower reliability compared to ICP. Decision Quality | positive | psychological plausibility and reliability of affective steering effects |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The same mechanisms that improve rationality also amplify sensitivity to affective interventions. Decision Quality | mixed | interaction between reasoning-enhancing interventions and responsiveness to affective steering |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Different steering methods trade off controllability against human-aligned behavior: ICP is more controllable but less human-aligned, while RLS is more human-aligned but less controllable/reliable. Governance And Regulation | mixed | trade-off between controllability and human-aligned behavior under different steering methods |
Reading fidelity
high
Study strength
medium
|
not reported
|
| There is a tension between reasoning (improving rationality) and affective steering, which has implications for both using LLMs as human simulators and for the safe deployment of LLM-based decision systems. Ai Safety And Ethics | mixed | safety/appropriateness of LLM deployment and fidelity as human simulators |
Reading fidelity
high
Study strength
speculative
|
not reported
|