3 cumulative citations
View corpus contextLarge language models inherit and magnify human inventory biases in dynamic ordering: GPT-4 overthinks and chases demand, producing larger errors, while the lean GPT-4o delivers near-optimal orders; providing the optimal formula does not remove the bias, implicating model architecture rather than knowledge gaps.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Problem definition: Although large language models (LLMs) are increasingly integrated into business decision making, their potential to replicate and even amplify human cognitive biases cautions a significant, yet not well-understood, risk. This is particularly critical in high-stakes operational contexts like supply chain management. To address this, we investigate the decision-making patterns of leading LLMs using the canonical newsvendor problem in a dynamic setting, aiming to identify the nature and origins of their cognitive biases. Methodology/results: Through dynamic, multi-round experiments with GPT-4, GPT-4o, and LLaMA-8B, we tested for five established decision biases. We found that LLMs consistently replicated the classic ``Too Low/Too High'' ordering bias and significantly amplified other tendencies like demand-chasing behavior compared to human benchmarks. Our analysis uncovered a ``paradox of intelligence'': the more sophisticated GPT-4 demonstrated the greatest irrationality through overthinking, while the efficiency-optimized GPT-4o performed near-optimally. Because these biases persist even when optimal formulas are provided, we conclude they stem from architectural constraints rather than knowledge gaps. Managerial implications: First, managers should select models based on the specific task, as our results show that efficiency-optimized models can outperform more complex ones on certain optimization problems. Second, the significant amplification of bias by LLMs highlights the urgent need for robust human-in-the-loop oversight in high-stakes decisions to prevent costly errors. Third, our findings suggest that designing structured, rule-based prompts is a practical and effective strategy for managers to constrain models' heuristic tendencies and improve the reliability of AI-assisted decisions.
Summary
Main Finding
LLMs reproduce and often amplify well-known human decision biases in a canonical inventory task (the newsvendor problem). In dynamic, multi-round experiments, GPT-4, GPT-4o, and LLaMA-8B all exhibited the classic “Too Low/Too High” ordering bias and strong demand-chasing; deviations persisted in risk-neutral settings and even when optimal formulas were provided. Notably, the more computationally advanced GPT-4 showed the largest departures from optimality (a “paradox of intelligence” caused by overthinking), whereas the efficiency-optimized GPT-4o performed near-optimally.
Key Points
- Five biases tested: (1) systematic ordering bias (“Too Low/Too High”), (2) presentation-order effect, (3) persistence of bias in risk-neutral settings, (4) demand-chasing (overreaction to recent demand), (5) constrained learning from feedback.
- All models replicated the classic ordering bias; some deviations were substantially larger than human benchmarks (up to ≈70% greater deviation reported for GPT-4).
- Biases persisted in risk-neutral scenarios → they are unlikely to be driven by modeled risk preferences and instead reflect information-processing/architectural constraints.
- Presentation-order effects: sequence of margin scenarios affected subsequent ordering behavior, producing path dependence similar to human anchoring.
- Demand-chasing: LLMs overweight recent demand realizations and adjust orders disproportionately; in several cases this amplification exceeded human tendency.
- Learning constraints: feedback produced adjustments but did not eliminate persistent deviations; learning plateaued at suboptimal levels.
- Paradox of intelligence: a more sophisticated model (GPT-4) exhibited greater irrationality through complex internal reasoning, while a model optimized for efficiency (GPT-4o) behaved closer to the normative optimum.
- Providing explicit analytical formulas (optimal-order solution) did not fully correct biases, implying limits arise from model architecture/processing rather than lack of knowledge.
Data & Methods
- Experimental framework: the canonical newsvendor problem implemented in dynamic, multi-round settings with feedback—designed to mirror operational, iterative decision contexts rather than single-shot tasks.
- Models tested: GPT-4, GPT-4o, and LLaMA-8B.
- Treatments and manipulations included: varied demand distributions, margin scenarios (high vs low margin), presentation-order sequencing, risk-neutral scenarios, multi-period demand realizations, and conditions with/without explicit optimal-ordering formulas.
- Evaluation metrics: deviations between model-chosen order quantities and analytical optimal order; measures of sensitivity to recent demand (demand-chasing); persistence of bias across rounds; comparisons to established human behavioral benchmarks from decades of newsvendor studies.
- Key empirical findings: consistent replication of five target biases across architectures, quantitatively larger deviations for GPT-4 in several cases, near-optimal performance by GPT-4o, and persistence of biases even when the optimal formula was provided.
Implications for AI Economics
- Model selection matters for economic tasks: greater architectural complexity does not guarantee better normative decision quality; efficiency-optimized or simpler models can outperform larger models on some optimization tasks.
- Risk of bias amplification: deploying LLMs in operational decision-making (inventory, forecasting, procurement) can magnify human-like biases, generating higher economic costs than human-only decision processes unless mitigated.
- Human-in-the-loop is essential: robust oversight, validation, and intervention mechanisms are needed in high-stakes supply chain and operational settings to detect and correct amplified biases.
- Prompt and system design: structured, rule-based prompts and constrained decision interfaces (explicit formula injection, checklist-style procedures, or restricted output formats) are practical mitigations to constrain heuristic tendencies and improve reliability.
- Evaluation standards: adopt canonical, analytically grounded benchmarks (like the newsvendor) and dynamic, multi-round tests when certifying LLMs for operational use—single-shot tests can miss path-dependent and time-evolving biases.
- Economic modeling and policy: incorporate LLM-induced behavioral distortions into cost-benefit analyses of AI deployment (expected error amplification, potential loss from demand-chasing, need for oversight), and consider regulatory/audit requirements for AI decision systems in supply chains.
- Research directions: investigate architectural causes of “overthinking” and bias amplification, expand experiments across more LLM families and real-world operational contexts, and develop debiasing techniques that target processing and reasoning pathways rather than only knowledge augmentation.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLMs consistently replicated the classic 'Too Low/Too High' ordering bias in the dynamic newsvendor task. Decision Quality | negative | ordering bias (Too Low/Too High orders) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLMs significantly amplified demand-chasing behavior compared to human benchmarks. Decision Quality | negative | demand-chasing behavior (tendency to chase recent demand when ordering) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| GPT-4 demonstrated the greatest irrationality (the paper's 'paradox of intelligence') by overthinking the problem and performing worse than less sophisticated models. Decision Quality | negative | rationality/optimality of ordering decisions (degree of departure from optimal order policy) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The efficiency-optimized model GPT-4o performed near-optimally on the optimization task, outperforming the more 'sophisticated' GPT-4 on this problem. Decision Quality | positive | optimality of ordering decisions / performance on newsvendor optimization |
Reading fidelity
high
Study strength
medium
|
not reported
|
| These biases persist even when optimal formulas are provided to the models, implying the biases stem from architectural constraints rather than knowledge gaps. Ai Safety And Ethics | negative | persistence of biased decisions despite provision of optimal formulas |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Managers should select models based on the specific task, because efficiency-optimized models can outperform more complex models on certain optimization problems. Organizational Efficiency | positive | model selection effectiveness for task-specific performance |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The significant amplification of bias by LLMs highlights the urgent need for robust human-in-the-loop oversight in high-stakes decisions to prevent costly errors. Organizational Efficiency | negative | risk of costly decision errors without human oversight |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Designing structured, rule-based prompts is a practical and effective strategy for managers to constrain models' heuristic tendencies and improve the reliability of AI-assisted decisions. Organizational Efficiency | positive | reliability of AI-assisted decisions under structured, rule-based prompting |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The study tested for five established decision biases in a dynamic, multi-round newsvendor setting using three leading LLMs (GPT-4, GPT-4o, and LLaMA-8B). Other | null_result | presence/absence of five established decision biases |
Reading fidelity
high
Study strength
high
|
not reported
|