The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models inherit and magnify human inventory biases in dynamic ordering: GPT-4 overthinks and chases demand, producing larger errors, while the lean GPT-4o delivers near-optimal orders; providing the optimal formula does not remove the bias, implicating model architecture rather than knowledge gaps.

Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
Jifei Liu, Zhi Chen, Yuanguang Zhong · December 14, 2025
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jifei Liu unresolved corpus identity
  2. Zhi Chen unresolved corpus identity
  3. Yuanguang Zhong unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jifei Liu provider ID
  2. Zhi Chen provider ID
  3. Yuanguang Zhong provider ID
In dynamic newsvendor experiments, leading LLMs systematically replicate and often amplify human ordering biases—GPT-4 exhibits pronounced ‘overthinking’ and demand-chasing while the efficiency-optimized GPT-4o performs near-optimally—and these biases persist even when optimal formulas are provided.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Problem definition: Although large language models (LLMs) are increasingly integrated into business decision making, their potential to replicate and even amplify human cognitive biases cautions a significant, yet not well-understood, risk. This is particularly critical in high-stakes operational contexts like supply chain management. To address this, we investigate the decision-making patterns of leading LLMs using the canonical newsvendor problem in a dynamic setting, aiming to identify the nature and origins of their cognitive biases. Methodology/results: Through dynamic, multi-round experiments with GPT-4, GPT-4o, and LLaMA-8B, we tested for five established decision biases. We found that LLMs consistently replicated the classic ``Too Low/Too High'' ordering bias and significantly amplified other tendencies like demand-chasing behavior compared to human benchmarks. Our analysis uncovered a ``paradox of intelligence'': the more sophisticated GPT-4 demonstrated the greatest irrationality through overthinking, while the efficiency-optimized GPT-4o performed near-optimally. Because these biases persist even when optimal formulas are provided, we conclude they stem from architectural constraints rather than knowledge gaps. Managerial implications: First, managers should select models based on the specific task, as our results show that efficiency-optimized models can outperform more complex ones on certain optimization problems. Second, the significant amplification of bias by LLMs highlights the urgent need for robust human-in-the-loop oversight in high-stakes decisions to prevent costly errors. Third, our findings suggest that designing structured, rule-based prompts is a practical and effective strategy for managers to constrain models' heuristic tendencies and improve the reliability of AI-assisted decisions.

Summary

Main Finding

LLMs reproduce and often amplify well-known human decision biases in a canonical inventory task (the newsvendor problem). In dynamic, multi-round experiments, GPT-4, GPT-4o, and LLaMA-8B all exhibited the classic “Too Low/Too High” ordering bias and strong demand-chasing; deviations persisted in risk-neutral settings and even when optimal formulas were provided. Notably, the more computationally advanced GPT-4 showed the largest departures from optimality (a “paradox of intelligence” caused by overthinking), whereas the efficiency-optimized GPT-4o performed near-optimally.

Key Points

  • Five biases tested: (1) systematic ordering bias (“Too Low/Too High”), (2) presentation-order effect, (3) persistence of bias in risk-neutral settings, (4) demand-chasing (overreaction to recent demand), (5) constrained learning from feedback.
  • All models replicated the classic ordering bias; some deviations were substantially larger than human benchmarks (up to ≈70% greater deviation reported for GPT-4).
  • Biases persisted in risk-neutral scenarios → they are unlikely to be driven by modeled risk preferences and instead reflect information-processing/architectural constraints.
  • Presentation-order effects: sequence of margin scenarios affected subsequent ordering behavior, producing path dependence similar to human anchoring.
  • Demand-chasing: LLMs overweight recent demand realizations and adjust orders disproportionately; in several cases this amplification exceeded human tendency.
  • Learning constraints: feedback produced adjustments but did not eliminate persistent deviations; learning plateaued at suboptimal levels.
  • Paradox of intelligence: a more sophisticated model (GPT-4) exhibited greater irrationality through complex internal reasoning, while a model optimized for efficiency (GPT-4o) behaved closer to the normative optimum.
  • Providing explicit analytical formulas (optimal-order solution) did not fully correct biases, implying limits arise from model architecture/processing rather than lack of knowledge.

Data & Methods

  • Experimental framework: the canonical newsvendor problem implemented in dynamic, multi-round settings with feedback—designed to mirror operational, iterative decision contexts rather than single-shot tasks.
  • Models tested: GPT-4, GPT-4o, and LLaMA-8B.
  • Treatments and manipulations included: varied demand distributions, margin scenarios (high vs low margin), presentation-order sequencing, risk-neutral scenarios, multi-period demand realizations, and conditions with/without explicit optimal-ordering formulas.
  • Evaluation metrics: deviations between model-chosen order quantities and analytical optimal order; measures of sensitivity to recent demand (demand-chasing); persistence of bias across rounds; comparisons to established human behavioral benchmarks from decades of newsvendor studies.
  • Key empirical findings: consistent replication of five target biases across architectures, quantitatively larger deviations for GPT-4 in several cases, near-optimal performance by GPT-4o, and persistence of biases even when the optimal formula was provided.

Implications for AI Economics

  • Model selection matters for economic tasks: greater architectural complexity does not guarantee better normative decision quality; efficiency-optimized or simpler models can outperform larger models on some optimization tasks.
  • Risk of bias amplification: deploying LLMs in operational decision-making (inventory, forecasting, procurement) can magnify human-like biases, generating higher economic costs than human-only decision processes unless mitigated.
  • Human-in-the-loop is essential: robust oversight, validation, and intervention mechanisms are needed in high-stakes supply chain and operational settings to detect and correct amplified biases.
  • Prompt and system design: structured, rule-based prompts and constrained decision interfaces (explicit formula injection, checklist-style procedures, or restricted output formats) are practical mitigations to constrain heuristic tendencies and improve reliability.
  • Evaluation standards: adopt canonical, analytically grounded benchmarks (like the newsvendor) and dynamic, multi-round tests when certifying LLMs for operational use—single-shot tests can miss path-dependent and time-evolving biases.
  • Economic modeling and policy: incorporate LLM-induced behavioral distortions into cost-benefit analyses of AI deployment (expected error amplification, potential loss from demand-chasing, need for oversight), and consider regulatory/audit requirements for AI decision systems in supply chains.
  • Research directions: investigate architectural causes of “overthinking” and bias amplification, expand experiments across more LLM families and real-world operational contexts, and develop debiasing techniques that target processing and reasoning pathways rather than only knowledge augmentation.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper uses controlled experiments and interventions (e.g., providing optimal formulas) that credibly show persistent, model-specific biases, which supports non-trivial causal interpretation about model behavior; however, evidence is limited to a simulated task, a small set of models and unspecified sample sizes/settings, and lacks field validation in real operational environments. Methods Rigormedium — The study employs a relevant dynamic task, multi-round design, tests for five canonical biases, and conducts direct interventions to probe mechanisms, but it appears to omit key reproducibility details (trial counts, prompt/temperature settings, random seeds), lacks pre-registration or robustness checks across a broader set of models and environments, and does not include real-world deployment evidence. SampleSimulated dynamic multi-round newsvendor experiments run on three LLMs (GPT-4, GPT-4o, LLaMA-8B) with tests for five established decision biases and comparisons to a human benchmark; specific number of rounds/replications and exact prompt/temperature configurations are not reported in the summary. Themeshuman_ai_collab org_design IdentificationControlled, multi-round behavioral experiments with three LLMs (GPT-4, GPT-4o, LLaMA-8B) on a simulated dynamic newsvendor task, plus intervention tests (providing the optimal ordering formula and structured/rule-based prompts) and comparisons to human benchmark behavior to infer whether biases arise from knowledge gaps versus model architecture/heuristics. GeneralizabilityResults derived from a stylized, simulated newsvendor task may not generalize to full-scale, real-world supply chain decisions with richer information and interactions., Only three model variants tested — findings may not hold across other LLM families, model sizes, or future model updates., Prompt wording, decoding settings, and system instructions (not fully specified) can materially affect behavior, limiting reproducibility., Human benchmark details are not fully described, so comparisons may depend on benchmark selection and experimental framing., Short-run laboratory-style experiments may not capture long-run adaptation, human oversight, or organizational processes that mitigate bias.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLMs consistently replicated the classic 'Too Low/Too High' ordering bias in the dynamic newsvendor task. Decision Quality negative ordering bias (Too Low/Too High orders)
Reading fidelity high
Study strength medium
not reported
0.18
LLMs significantly amplified demand-chasing behavior compared to human benchmarks. Decision Quality negative demand-chasing behavior (tendency to chase recent demand when ordering)
Reading fidelity high
Study strength medium
not reported
0.18
GPT-4 demonstrated the greatest irrationality (the paper's 'paradox of intelligence') by overthinking the problem and performing worse than less sophisticated models. Decision Quality negative rationality/optimality of ordering decisions (degree of departure from optimal order policy)
Reading fidelity high
Study strength medium
not reported
0.18
The efficiency-optimized model GPT-4o performed near-optimally on the optimization task, outperforming the more 'sophisticated' GPT-4 on this problem. Decision Quality positive optimality of ordering decisions / performance on newsvendor optimization
Reading fidelity high
Study strength medium
not reported
0.18
These biases persist even when optimal formulas are provided to the models, implying the biases stem from architectural constraints rather than knowledge gaps. Ai Safety And Ethics negative persistence of biased decisions despite provision of optimal formulas
Reading fidelity high
Study strength medium
not reported
0.18
Managers should select models based on the specific task, because efficiency-optimized models can outperform more complex models on certain optimization problems. Organizational Efficiency positive model selection effectiveness for task-specific performance
Reading fidelity high
Study strength speculative
not reported
0.03
The significant amplification of bias by LLMs highlights the urgent need for robust human-in-the-loop oversight in high-stakes decisions to prevent costly errors. Organizational Efficiency negative risk of costly decision errors without human oversight
Reading fidelity high
Study strength speculative
not reported
0.03
Designing structured, rule-based prompts is a practical and effective strategy for managers to constrain models' heuristic tendencies and improve the reliability of AI-assisted decisions. Organizational Efficiency positive reliability of AI-assisted decisions under structured, rule-based prompting
Reading fidelity high
Study strength speculative
not reported
0.03
The study tested for five established decision biases in a dynamic, multi-round newsvendor setting using three leading LLMs (GPT-4, GPT-4o, and LLaMA-8B). Other null_result presence/absence of five established decision biases
Reading fidelity high
Study strength high
not reported
0.3

Notes