The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Enterprises that default to the largest LLM risk unnecessary emissions without quality gains: a 520-task supply-chain benchmark finds smaller or differently architected models can match accuracy while producing orders-of-magnitude less per-task carbon, and a carbon-aware procurement framework plus calibrated routing cuts estimated inference carbon by roughly 98% while preserving or improving decision quality.

Toward Sustainable AI Deployment: A Carbon-Aware Decision Framework for Enterprise Supply Chain Systems
Haoran Yu, Lifei Liu, Danping Zhang · September 14, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Haoran Yu unresolved corpus identity
  2. Lifei Liu unresolved corpus identity
  3. Danping Zhang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Hao-Ran Yu unresolved corpus identity
  2. Li-Fei Liu unresolved corpus identity
  3. Dan-Ping Zhang unresolved corpus identity
Benchmarking six LLMs on 520 supply-chain tasks shows substantial variation in decision quality and estimated inference-related carbon, demonstrates that larger parameter counts do not reliably predict better task performance, and proposes a Carbon-Aware AI Procurement Framework (CAAPF) with a GreenRoute routing proof-of-concept that achieves large estimated carbon savings while meeting quality thresholds.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Enterprises deploying AI for supply chain decisions commonly default to the largest available language model, a procurement heuristic that neglects both empirical performance and environmental cost. We benchmark six large language models across 520 supply chain tasks, simultaneously measuring decision quality and estimated generation-related operational carbon. Drawing on the Technology-Organization-Environment (TOE) framework, we develop a Carbon-Aware AI Procurement Framework (CAAPF), a Green IS design artifact that operationalizes sustainable AI governance for enterprise procurement. Within this bounded sample, quality spans 0.497-0.723, and the models with the largest disclosed parameter totals do not achieve the highest scores. The design does not isolate size, provider, architecture, or benchmark-construction effects. A category-by-tier calibrated GreenRoute proof of concept reaches 0.733 mean out-of-sample quality at an estimated 0.402 gCO2/task. Static Haiku reaches 0.699 at 0.022 gCO2/task, while Sonnet reaches 0.723 at 0.401 gCO2/task, demonstrating that the preferred strategy depends on the organization's quality requirement. Our "benchmark first, select green" principle suggests that environmental responsibility and decision quality can be mutually reinforcing, contributing to sustainable digital infrastructure governance aligned with SDG 12 and SDG 13.

Summary

Main Finding

Benchmarking six LLMs on 520 supply-chain tasks shows that larger nominal model size is an unreliable proxy for decision quality, and that substantial variation in operational (inference) carbon exists across models. A TOE-grounded Carbon-Aware AI Procurement Framework (CAAPF) — operationalized as a category×tier calibrated routing policy (GreenRoute) — can deliver materially higher quality per unit carbon than naïve “choose the biggest model” procurement. In this sample, the Pareto frontier is formed by Claude Haiku (low carbon, near-top quality) and Claude Sonnet (highest quality, moderate carbon); a simple low-carbon static choice or a calibrated routing policy is often preferable depending on an organization’s quality threshold.

Key Points

  • Models evaluated (aggregate score on 520 tasks):
    • Claude Sonnet 4.6: score 0.723; Tier‑1 GT 78.6%
    • Claude Haiku 4.5: score 0.699; Tier‑1 GT 75.3%
    • Mistral Large 3: 0.613
    • Llama 3.3 70B: 0.528
    • Llama 4 Scout: 0.521
    • DeepSeek R1: 0.497
  • Estimated generation-related per-task carbon (study assumptions; ±50% uncertainty):
    • Haiku: 0.022 gCO2/task (312 tokens/task)
    • Scout: 0.023 gCO2/task (387 tokens/task)
    • Sonnet: 0.401 gCO2/task (445 tokens/task)
    • Llama 3.3 70B: 0.468 gCO2/task (478 tokens/task)
    • Mistral Large 3: 3.427 gCO2/task (623 tokens/task)
    • DeepSeek R1: 17.082 gCO2/task (2,847 tokens/task)
  • Verbosity amplification effect: per-token carbon rates understate workload-level emissions when models generate vastly different output lengths (DeepSeek generates ~9× more tokens than Haiku and produces far higher per-task emissions despite lower accuracy).
  • SC-domain knowledge (30 “trap” tasks): Claude models substantially outperform others (Sonnet 0.893, Haiku 0.861), indicating domain-relevant differences not captured by parameter counts.
  • CAAPF principle: “benchmark first, select green.” Steps: define task portfolio & thresholds (τ), benchmark candidates, measure quality/carbon/cost/latency, filter by admissibility, select lowest-carbon admissible model, monitor & re-evaluate.
  • GreenRoute (proof of concept routing calibrated per category×tier with τ=0.65):
    • Out-of-sample mean quality 0.733 (SD 0.004)
    • Mean estimated carbon 0.402 gCO2/task (SD 0.124)
    • Estimated savings vs. always-DeepSeek: 97.6%
    • Choice logic: route only when task‑level heterogeneity improves the feasible frontier; otherwise prefer static lowest-carbon passing model.
  • Propositions:
    • P1: Model size is not a reliable procurement proxy when jointly considering quality and operational carbon.
    • P2: Preferred selection strategy depends on organization-specific quality/risk thresholds; routing is not always beneficial.
    • P3: A TOE-based (Technology‑Organization‑Environment) process creates an auditable procurement record incorporating environmental pressures without making carbon the sole criterion.

Data & Methods

  • Benchmark:
    • 520 supply-chain tasks across 6 categories: Demand Forecasting (92, incl. 33 M5 tasks), Vehicle Routing (84), Inventory Optimization (96), Supplier Risk Assessment (88), Order Fulfillment (80), Demand Classification (80).
    • Three difficulty tiers: Tier 1 (formula/deterministic ~35%), Tier 2 (multi-factor ~40%), Tier 3 (strategic judgment ~25%).
    • 30 SC‑trap tasks specifically probing supply-chain domain knowledge.
    • Tasks authored and validated internally (three-stage validation), but not validated by external practitioners.
  • Models: six availability‑based LLMs from four providers (sample concentrated among a few vendors; Anthropic models lacked disclosed parameter counts).
  • Prompting & evaluation:
    • Single prompt protocol, model temperature = 0, output caps 1,024–2,048 tokens.
    • Dual LLM-as-judge evaluators: Claude Opus 4 and Qwen3-235B; scored numeric accuracy, reasoning quality, and domain knowledge (1–10 scales), averaged and normalized.
    • Tier-1 tasks validated against deterministic ground truth (exact match within 5%) to control for judge bias.
  • Carbon estimation:
    • Model-specific gCO2/MTok coefficients computed from serving hardware, GPU power, throughput, PUE (1.1–1.2), and a US‑East grid intensity assumption (0.38 kgCO2/kWh); estimates carry ±50% uncertainty.
    • Per-task carbon = gCO2/MTok × mean output tokens / 1e6. Input/prefill tokens not included (so values reflect output-generation carbon only).
  • GreenRoute:
    • Calibrated per category×tier (18 cells); cell-level τ = 0.65 used to select lowest-carbon model that meets threshold (else highest-quality).
    • Evaluation: 20 repeated stratified 50/50 split-half trials (seed 42); reported out‑of‑sample metrics.

Limitations noted by authors: - Sample not representative of global LLM market (vendor concentration and disclosure gaps). - Carbon estimates are engineering assumptions, approximate and output-only (±50%). - Single prompting protocol; did not explore demonstrations, prompt engineering, or temperature effects. - Judges are LLMs (possible evaluator biases) and provider defaults differ slightly. - Confounding across provider, architecture, training data; causal claims about why differences arise are limited. - GreenRoute proof-of-concept does not evaluate production metadata classification costs, routing overhead, or real-world governance integration.

Implications for AI Economics

  • Procurement heuristics: The common “choose the largest model” rule can be economically and environmentally suboptimal. Organizations should benchmark domain-specific task quality and account for operational carbon when procuring models.
  • Internalizing energy externalities: Including estimated inference carbon in procurement decisions (and reporting) creates incentives for providers to optimize inference efficiency and for buyers to choose lower-carbon models when acceptable quality is achieved.
  • Cost-quality-carbon tradeoffs: Organizations with moderate quality requirements can often pick a low-carbon static model (e.g., Haiku) and obtain large emissions reductions for little quality loss. High-quality requirements may justify higher-carbon choices (e.g., Sonnet) or calibrated routing that mixes models across task classes.
  • Routing economics: Task-level routing can raise system complexity and overhead; it is only justified if it improves the feasible quality–carbon frontier given organization-specific thresholds. The paper’s CAAPF operationalizes the decision rule (benchmark first; route only if needed).
  • Market signals & provider incentives: Empirical procurement that rewards low-carbon, high-quality inference will pressure providers to disclose per-inference energy metrics, invest in more efficient architectures, or offer differentiated pricing for efficient endpoints.
  • Policy & governance: TOE-informed, auditable procurement records that include carbon measurements align with SDG 12/13 and enable compliance reporting, internal carbon accounting, and regulated procurement standards.
  • Research & measurement needs: Better standardized per-request energy/carbon reporting from providers, end-to-end serving measurements (including input/prefill tokens and infrastructure variability), and evaluations across more diverse and representative models and prompting protocols are required for robust economic policy design.

Bottom line: For enterprise supply-chain AI, benchmarking domain-specific decision quality together with operational carbon yields better procurement choices than naive size-based rules. CAAPF offers a practical, auditable decision sequence — “benchmark first, select green” — that firms can adopt to align AI deployments with economic and environmental objectives.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic benchmarking across 520 domain-specific tasks with a clear evaluation protocol and an implemented procurement/routing proof-of-concept, supporting descriptive claims about quality–carbon tradeoffs. However, key limitations (non-representative model sample, LLM-based judges, large uncertainty in carbon estimates, output-only energy accounting, and no external/human validation) reduce the confidence in generalizing the empirical magnitudes. Methods Rigormedium — Strengths: well-structured task taxonomy, multi-stage task validation, fixed prompting and temperature controls, deterministic checks for Tier-1 tasks, repeated split-half trials for GreenRoute, and transparent reporting of assumptions. Weaknesses: small and availability-biased model sample, reliance on LLM-as-judge rather than independent human raters, large ±50% uncertainty in carbon engineering estimates and omission of end-to-end serving tokens, and no external validation of task realism or production routing overheads. SampleSix LLMs from four providers (Anthropic Claude Sonnet & Haiku, Mistral Large 3, Meta Llama variants, DeepSeek R1) evaluated on 520 supply-chain tasks across 6 categories (Demand Forecasting, Vehicle Routing, Inventory Optimization, Supplier Risk, Order Fulfillment, Demand Classification) and 3 difficulty tiers; two LLM-based judges (Claude Opus 4 and Qwen3-235B) scored responses; estimated generation-related carbon per task computed from model-specific gCO2/Mtok engineering coefficients (±50% uncertainty) multiplied by mean output tokens; GreenRoute routing calibrated on 18 category×tier cells with τ=0.65 using repeated stratified split-half trials. Themesgovernance adoption GeneralizabilityModel sample is availability-biased and over-represents certain providers; results may not generalize to other LLMs or future model families., Carbon estimates are engineering approximations with ±50% uncertainty and measure only output-generation energy (input/prefill and end-to-end serving energy omitted)., Evaluation uses fixed prompting and temperature=0, so findings do not account for prompt engineering, few-shot examples, or alternative decoding strategies that can affect quality and verbosity., Quality adjudication uses LLM-as-judge rather than independent human experts, which may introduce evaluator bias or circularity., Benchmark tasks were author-validated but not externally validated by practicing supply-chain professionals; real-world organizational context, latency, integration costs, and governance constraints are not fully captured., Assumed cloud region and PUE (US-East intensity, PUE 1.1–1.2) limit transferability to other geographies or infrastructure setups.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Across the six evaluated language models, quality on 520 supply-chain tasks ranged from 0.497 to 0.723, and the models with the largest disclosed parameter totals were not the highest-scoring models. Output Quality negative Aggregate supply-chain decision quality
Reading fidelity high
Study strength medium
n=520
Quality range: 0.497–0.723
0.18
Claude Haiku 4.5 achieved nearly the quality of Claude Sonnet 4.6 while producing substantially less estimated generation-related carbon per task. Organizational Efficiency mixed Decision quality and estimated generation-related carbon per task
Reading fidelity high
Study strength medium
n=520
Haiku: 0.699 quality and 0.022 gCO2/task; Sonnet: 0.723 quality and 0.401 gCO2/task
0.18
DeepSeek R1 generated much longer responses than Haiku and had substantially higher estimated output-generation carbon per task despite lower benchmark quality. Organizational Efficiency negative Estimated generation-related carbon per task and decision quality
Reading fidelity high
Study strength medium
n=520
780.4× Haiku's estimated carbon per task
0.18
For DeepSeek R1's formula tasks, longer responses were monotonically associated with lower accuracy. Output Quality negative Formula-task accuracy
Reading fidelity high
Study strength medium
n=91
Accuracy declined from 18.2% to 4.5% across response-length quartiles
0.18
The gap between Claude Sonnet 4.6 and DeepSeek R1 increased as task difficulty rose from Tier 1 to Tier 3. Output Quality negative Decision quality by task difficulty tier
Reading fidelity high
Study strength medium
Sonnet–DeepSeek quality gap: 0.192 in Tier 1 and 0.257 in Tier 3
0.18
Claude models performed best on the benchmark's supply-chain domain-knowledge checks, while DeepSeek R1 scored lower. Output Quality positive Supply-chain domain-knowledge task score
Reading fidelity high
Study strength medium
n=30
Sonnet: 0.893; Haiku: 0.861; DeepSeek R1: 0.687
0.18
The GreenRoute proof-of-concept routing policy achieved higher mean out-of-sample quality than always using DeepSeek R1 while reducing estimated per-task carbon by 97.6%. Organizational Efficiency positive Out-of-sample decision quality and estimated carbon savings
Reading fidelity high
Study strength medium
n=20
0.733 mean quality; 0.402 gCO2/task; 97.6% estimated carbon savings
0.18
Across GreenRoute trials, 82.2% of held-out cell–trial means met the calibration quality threshold of 0.65. Output Quality positive Proportion of held-out category-by-tier means meeting the quality threshold
Reading fidelity high
Study strength medium
n=360
296 of 360 means, or 82.2%, met τ = 0.65
0.18
The paper's proposed CAAPF framework recommends selecting the lowest-carbon model among candidates that satisfy organization-specific quality and governance constraints, rather than selecting a model solely by size. Governance And Regulation positive Sustainable and auditable AI procurement decision process
Reading fidelity high
Study strength speculative
not reported
0.03

Notes