The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A 0–100 'sustainability score' that blends carbon intensity, utilisation and a hardware bonus is proposed to nudge researchers toward greener cluster usage, with a 12‑week randomized crossover trial on a university ML cluster planned to test its behavioural impact; results are pending.

Scoring and Gamification to Encourage Sustainable Use of Compute Clusters
Maximilian MacDonald, Chris McCaig, Sean MacAvaney, Matthew Barr, Lauritz Thamsen · August 19, 2026
arxiv descriptive low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Maximilian MacDonald unresolved corpus identity
  2. Chris McCaig unresolved corpus identity
  3. Sean MacAvaney unresolved corpus identity
  4. Matthew Barr unresolved corpus identity
  5. Lauritz Thamsen unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Maximilian MacDonald provider ID
  2. Christopher McCaig provider ID
  3. Sean MacAvaney provider ID
  4. Matthew Barr provider ID
  5. L. Thamsen provider ID
The paper proposes a gamified composite 0–100 sustainability score (combining average carbon intensity, resource utilisation, and an underutilised-hardware bonus) and outlines a 12-week within-subjects study on a university ML cluster to test whether dashboard feedback encourages more sustainable compute behaviour.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The environmental cost of computing continues to grow, yet behaviour change remains limited. We present a composite sustainability score integrating average carbon intensity, resource utilisation, and embodied emissions into a single 0-100 metric designed for gamified feedback. Each component rewards a different dimension of sustainable behaviour: carbon-aware workload shifting, high resource utilisation, and selecting hardware that is commonly underutilised. This scoring system is built into an existing cluster management interface and underpins three dashboard conditions: raw metrics, composite score, and a gamified tree visualisation, which we are planning to evaluate in a 12-week within-subjects study with approximately 35 researchers. Furthermore, we open the discussion on the challenge of defining computational work `goodness' in the context of sustainability scores.

Summary

Main Finding

The authors design a gamifiable composite sustainability score for compute-cluster users that combines average carbon intensity, resource utilisation, and a hardware “bonus” for using underutilised machines into a single 0–100 metric. They embed this score into three dashboard conditions (raw metrics, composite score, and a gamified forest visualization) and plan a 12‑week within-subjects study (≈35 researchers) to test whether gamified feedback changes attitudes and resource-usage behaviour.

Key Points

  • Motivation: Numerical energy/carbon metrics alone are not action-guiding; multi‑dimensional sustainability (operational carbon, embodied carbon, utilisation) requires a balanced metric suitable for behavioural interventions.
  • Composite score components:
    • Utilisation (Uscore): average of GPU, CPU, GPU‑mem, CPU‑mem utilisation (normalized 0–1; capped at 1 for over‑use).
    • Hardware bonus (Hbonus): incentive for using underutilised GPU types computed inversely from rolling 1‑week cluster utilisation.
    • Average Carbon Intensity score (ACIscore): compares actual job timing to a feasible optimal contiguous window (1 week or twice job duration) and maps between best and worst-case ACI.
  • Default weights: wU = 0.4, wH = 0.4, wACI = 0.2 (sum = 1). Final score: Stotal = wU·Uscore + wH·Hbonus + wACI·ACIscore.
  • Gamification: composite score mapped to an interactive forest (trees = pods; leaf colour = hardware bonus; ground/gradients = utilisation; weather = carbon intensity) to leverage intrinsic/extrinsic motivation and visual feedback.
  • Study design: 12 weeks total; participants rotate through three dashboard conditions randomized in order, each shown for 3 weeks, plus 3 weeks of control; mixed methods — continuous behavioral logging and pre/post surveys, plus 10–15 qualitative interviews (reflexive thematic analysis).
  • Measurement details: integration into Launcher (cluster web app) on OKD Kubernetes; GPU power measured via NVIDIA SMI; CPU power estimated via linear model to TDP; resource metrics sampled every 5 minutes; grid intensity from UK Carbon Intensity API.
  • Open challenges acknowledged: defining “good” or useful computational work (including failed jobs), gaming the metric (under‑provisioning), fairness for users with high legitimate resource needs, and limits of CPU power estimation.

Data & Methods

  • Setting: mid‑sized ML research cluster at University of Glasgow; ~35 active users.
  • Intervention:
    • Conditions: (1) Raw metrics baseline (energy kWh, CO2e, average carbon intensity); (2) Composite sustainability score with breakdown; (3) Gamified forest interface visualizing the composite score and subcomponents.
    • Exposure schedule: randomized within-subject rotation; 3 weeks per dashboard + 3 weeks control (total 12 weeks).
  • Measurements:
    • Operational carbon: real-time regional grid intensity × estimated/measured hardware power.
    • Power measurement: GPU via NVML/NVIDIA SMI; CPU via TDP-based linear model (noted accuracy limits at low utilisation).
    • Resource utilisation: CPU, GPU, memory measured at node-level every 5 minutes.
    • Hardware bonus: cluster-wide rolling 1‑week utilisation per GPU type.
    • ACI calculation: for each job, find contiguous block of timestamps (≥ job length, bounded by 1 week or 2× job duration) with min/max average carbon intensity; score linearly between optimal and worst-case.
  • Evaluation data:
    • Behavioral logs: job submissions, allocations, utilisation, hardware selection, dashboard engagement.
    • Surveys: baseline and post-intervention on awareness, attitudes, perceived barriers, score interpretation.
    • Qualitative interviews: 10–15 semi-structured follow-ups analyzed with reflexive thematic analysis.
  • Planned analyses: within-subject comparisons of behaviour across conditions (e.g., timing shifts, utilisation changes, hardware choices), attitude changes, and qualitative exploration of interpretation and barriers.

Implications for AI Economics

  • Behavioral levers for compute externalities:
    • The paper operationalizes a low-cost, UI‑based intervention that can shift researcher behaviour (timing, resource provisioning, hardware selection) without changing pricing — a complement to price mechanisms (e.g., internal carbon price, chargeback).
    • Measuring behavioural elasticities (how much carbon/usage changes in response to score/gamification) would inform the effectiveness and cost-efficiency of non‑price interventions versus monetary incentives.
  • Accounting for embodied vs operational carbon:
    • Explicitly rewarding use of underutilised (often older) hardware can reduce embodied carbon pressure from accelerating hardware replacement. AI economists should consider embodied emissions when designing incentives and allocation policies; treating only operational margins can bias procurement decisions.
  • Productivity–emissions tradeoffs and welfare:
    • Interventions that encourage delaying jobs or increasing utilisation may affect researcher productivity (deadlines, iteration speed). Economic evaluation requires estimating the welfare tradeoff: reductions in emissions versus potential productivity loss or delayed output value.
    • Research contexts involve high uncertainty and exploratory work; metrics that penalize “low-value” or failed runs risk reducing socially valuable experimentation. Mechanism design must balance incentives with incentives for innovation.
  • Risks of perverse incentives and gaming:
    • The score can be gamed (under‑requesting resources, artificially delaying jobs) or create inequities (users with time‑sensitive/high‑compute tasks disadvantaged). AI economists designing allocation or pricing policies should simulate and measure gaming responses and distributional impacts.
  • Policy and institutional design opportunities:
    • Composite scores could be integrated into internal governance (soft quotas, dashboards, or tied to internal cost allocations) as a low-friction policy tool.
    • Such scores can inform procurement strategy (buying hardware that reduces total lifecycle emissions) and capacity planning by revealing which hardware is chronically underutilised.
  • Research suggestions for AI economists:
    • Run randomized controlled trials to estimate causal impacts of gamified scores on carbon emissions, utilisation, and productivity.
    • Quantify rebound effects (e.g., does making compute “appear cheaper” via reuse incentives increase total job volume?), and net emissions.
    • Explore hybrid interventions combining UI feedback with price signals (e.g., dynamic internal pricing tied to ACI) and evaluate cost-effectiveness.
    • Model welfare effects of discouraging failed or exploratory jobs and design credit mechanisms (e.g., “exploration budgets”) to preserve research innovation.
    • Extend generalizability analysis: different grid mixes, cloud vs on‑premise, diverse user populations, and alternative CPU power measurement methods.
  • Limitations to account for in economic analysis:
    • Measurement noise (CPU power model), local grid specifics (UK API used), and contextual constraints (deadlines, collaboration norms) limit external validity. Weighting of score components is subjective and may need calibration to local incentives and social objectives.

Summary recommendation: The proposed composite score and gamified feedback are promising, low-friction policy tools to influence researcher behaviour on clusters. For AI economics, the next step is rigorous causal evaluation of emissions reductions, productivity impacts, gaming behavior, and cost-effectiveness relative to pricing or allocation policies — with careful attention to embodied carbon, equity, and preserving exploratory research.

Assessment

Paper Typedescriptive Evidence Strengthlow — The manuscript describes design and a planned study but presents no empirical results; evidence will depend on a small (≈35) sample and short (12-week) deployment, so currently there is no demonstrated causal evidence. Methods Rigormedium — Design uses a sensible within-subject randomized crossover with counterbalancing and mixed methods (behavioural logs + surveys + interviews), and includes direct GPU metering and frequent utilisation sampling; limitations include small sample, short blocks (potential carryover), estimated CPU power via a simple TDP model, potential for gaming the score, and challenges attributing failed jobs. SamplePlanned deployment on a mid-sized ML research cluster at University of Glasgow with ~35 active researcher users; 12-week within-subjects crossover (three 3-week dashboard conditions + 3-week control) with continuous logging of job submissions, resource allocation, CPU/GPU/memory utilisation (5-minute intervals), GPU power via NVIDIA SMI, CPU power estimated from TDP models, cluster-wide GPU-type utilisation for hardware bonus, UK grid carbon intensity from National Grid API; baseline and post surveys and semi-structured interviews with ~10–15 participants. Themesadoption org_design IdentificationPlanned within-subjects randomized crossover: participants rotate through three dashboard conditions (raw metrics, composite score, gamified view) in counterbalanced order with 3-week blocks plus a 3-week control; behavioural logs and pre/post surveys enable comparison of behaviour and attitudes across conditions to estimate intervention effects. GeneralizabilitySmall, single-institution sample of ML researchers limits representativeness to other user populations (industry, other disciplines)., UK grid carbon intensity context may not generalize to regions with different grid mixes or temporal patterns., Mid-sized research cluster with specific hardware heterogeneity; findings may not transfer to large-scale production clouds or homogeneous clusters., Short study duration and 3-week condition blocks may not capture long-term behavioural change or seasonal effects., Scoring design assumptions (weights, power estimation models, hardware bonus) may not generalize and are subject to gaming or reconfiguration by users.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The paper presents a composite sustainability score integrating average carbon intensity, resource utilisation, and hardware lifecycle considerations into a single 0–100 metric intended for gamified feedback. Organizational Efficiency positive Sustainable compute-cluster usage
Reading fidelity high
Study strength speculative
not reported
0.03
The planned evaluation will test whether composite and gamified sustainability feedback encourages more sustainable computing practices. Organizational Efficiency positive Sustainable computing practices and user attitudes
Reading fidelity high
Study strength speculative
n=35
0.03
The planned study compares raw metrics, a composite sustainability score, and a gamified tree visualisation, with each condition presented for three weeks alongside a three-week control period. Organizational Efficiency positive Behavioral differences across dashboard conditions
Reading fidelity high
Study strength speculative
n=35
0.03
The proposed sustainability score assigns initial weights of 0.4 to resource utilisation, 0.4 to the hardware bonus, and 0.2 to average carbon intensity. Organizational Efficiency positive Composite sustainability score
Reading fidelity high
Study strength speculative
w_U=0.4, w_H=0.4, w_ACI=0.2
0.03
The resource-utilisation component averages GPU usage, CPU usage, GPU memory usage, and CPU memory usage, scoring each dimension linearly from 0 at zero utilisation to 1 at 100% utilisation. Organizational Efficiency positive Compute-resource utilisation efficiency
Reading fidelity high
Study strength speculative
not reported
0.03
The hardware-bonus component gives the greatest incentive to use the least-utilised GPU type in the cluster and no bonus to fully utilised GPU types. Task Allocation positive Selection of underutilised compute hardware
Reading fidelity high
Study strength speculative
not reported
0.03
The average-carbon-intensity component rewards scheduling jobs during relatively low-carbon periods within a constrained scheduling window. Task Allocation positive Carbon-aware workload timing
Reading fidelity high
Study strength speculative
not reported
0.03
Prior work cited in the paper found that gamified sustainability interventions in related domains can reduce energy consumption, including reductions of up to 20% in residential settings. Organizational Efficiency positive Energy consumption
Reading fidelity high
Study strength low
up to 20% reductions
0.09
A previously studied gamified garden visualisation reduced office users’ energy consumption by 0.5 standard deviations. Organizational Efficiency positive Office energy consumption
Reading fidelity high
Study strength low
0.5 standard deviations
0.09
Carbon-aware workload shifting has been reported to reduce workload carbon footprints by up to 70% in prior work using CASPER. Organizational Efficiency positive Workload carbon footprint
Reading fidelity high
Study strength low
up to 70%
0.09
The paper identifies a risk that poorly designed sustainability scores could encourage users to game the metric or create inequitable competition for users with legitimately high-resource workloads. Ai Safety And Ethics negative Fairness and validity of sustainability feedback
Reading fidelity high
Study strength speculative
not reported
0.03

Notes