The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Amsterdam’s evaluation finds no one-size-fits-all LLM for government: models that are more accurate tend to be costlier and more energy-intensive, while bias levels vary independently, meaning municipal procurement must balance quality, sustainability and fairness.

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Laurens Samson, Iva Gornishka, Gossa Lô, Yuki M. Asano, Sennay Ghebreab · August 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Laurens Samson unresolved corpus identity
  2. Iva Gornishka unresolved corpus identity
  3. Gossa Lô unresolved corpus identity
  4. Yuki M. Asano unresolved corpus identity
  5. Sennay Ghebreab unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Laurens Samson provider ID
  2. Iva Gornishka provider ID
  3. Gossa Lô provider ID
  4. Yuki M. Asano provider ID
  5. S. Ghebreab provider ID
A government-centred Dutch LLM benchmark shows no single model dominates across factuality, honesty, bias, energy and cost, forcing explicit trade-offs when selecting models for municipal use.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.

Summary

Main Finding

No single LLM is optimal for Dutch governmental use — measurable trade-offs are unavoidable. Higher-quality models (factuality/honesty) tend to incur greater energy use and financial cost, while social bias behaves largely independently of those trade-offs. Factuality and honesty are distinct properties: a model can answer correctly but still fail to acknowledge its limits.

Key Points

  • Framework developed with City of Amsterdam stakeholders (advisory board, practitioner interviews, and a 429-user chatbot survey) to translate organisational values into measurable evaluation dimensions.
  • Six evaluation dimensions prioritised for governmental deployment: factuality, honesty, social bias, energy consumption (sustainability), cost, and training-data transparency. Two practical use cases were emphasised: text simplification and summarisation.
  • Large-scale, Dutch-focused benchmark suite applied to 30+ multilingual and Dutch-specific models (including GEITje, Fietje, GPT‑NL and others).
  • Empirical findings:
    • No model dominates across all dimensions.
    • Quality (factuality/honesty) correlates positively with energy consumption and cost.
    • Social bias does not correlate strongly with either cost or energy; it requires separate measurement and mitigation.
    • High factuality does not imply high honesty — they must be measured separately.
  • Usability: results are translated into a user-friendly, non-technical leaderboard and public overview to support procurement and policy decisions.
  • Code and interactive overview released publicly (repository and web dashboard).

Data & Methods

  • Stakeholder input:
    • Advisory board: 9 invited experts (7 responded to the questionnaire). Process: open questionnaire → deep-dive sessions → iterative feedback.
    • Practitioner interviews: 18 staff across roles (data scientists, product owners, managers, policymakers).
    • End-user survey: 429 internal chatbot users (prioritised factuality/quality more highly than advisory board).
  • Benchmarks and operationalisation:
    • Factuality: Dutch or verified human-translated knowledge benchmarks; evaluations economised via tinyBenchmarks (≈100 curated examples per benchmark) to scale across many models.
    • Honesty: measured with HONESTCITYBENCH (developed alongside this project) to capture whether models acknowledge limitations/refuse when appropriate.
    • Social bias: used Dutch-adapted bias datasets (e.g., MBBQ, Burema) covering protected characteristics (age, origin, disability, gender).
    • Sustainability (energy): inference energy tracked with CodeCarbon to estimate environmental impact.
    • Cost: estimated from API pricing and self-hosting operational costs to reflect deployment budget impacts.
    • Training-data transparency: categorical/qualitative assessment of model license and available documentation on data provenance and collection practices.
  • Model coverage: over 30 Dutch-specific and multilingual generative models were evaluated, enabling cross-model comparisons.
  • Presentation: technical scores converted into interpretable categorical ratings (1–5 scales, coloured pills, and normalised bar charts) for non-technical stakeholders.

Implications for AI Economics

  • Multi-dimensional procurement decisions: Governments should avoid single-metric selection (e.g., highest accuracy). Procurement frameworks must operationalise explicit trade-offs (quality vs cost vs carbon).
  • Internalisation of externalities: Energy (carbon) costs correlate with higher-performing models — public buyers should account for environmental externalities in total-cost-of-ownership and may require carbon budget constraints or prefer energy-efficient alternatives for low-risk tasks.
  • Pricing and market structure:
    • Higher-quality models commanding higher inference costs could reinforce market concentration (providers able to fund large-scale models). Public procurement that values openness and training-data transparency may help sustain competition and local/national model development.
    • There is room for differentiated model choice by use case: lower-cost, lower-energy models may suffice for routine, non-critical tasks (e.g., basic drafting, summarisation), while high-criticality tasks require higher-quality models plus stronger oversight.
  • Bias requires separate regulation and monitoring: since social bias is not strongly tied to cost or energy, cost-based procurement alone won’t mitigate fairness risks. Regular bias audits and remediation investments are necessary.
  • Honesty and trustworthiness as economic factors: models that better acknowledge uncertainties reduce downstream verification costs and potential liability; quantifying honesty can inform expected verification overheads and public-trust trade-offs.
  • Support for local-language and open models: investing in Dutch-specific and transparent models can reduce dependency on expensive multilingual commercial offerings, potentially lowering both monetary and environmental costs while improving data-governance alignment (GDPR, copyright).
  • Policy recommendations for governments:
    • Adopt multi-criteria decision frameworks that explicitly weight factuality, honesty, bias, cost, energy, and data transparency.
    • Require reporting of inference energy and training-data provenance in procurement bids.
    • Fund or procure energy-efficient models and tools for retrieval-augmented systems (which may reduce reliance on large parametric models).
    • Mandate independent bias and honesty evaluations as part of deployment approvals.

If you want, I can (a) extract the concrete numeric results for specific models from the paper and present them in a compact table, or (b) draft a short procurement checklist based on the framework for municipal buyers. Which would be most useful?

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper presents a broad, systematic evaluation of 30+ LLMs across multiple dimensions (factuality, honesty, bias, energy, cost, training-data transparency) grounded in advisory-board values, practitioner interviews (n=18) and a chatbot-user survey (n=429). It combines established benchmarks, a newly developed honesty benchmark, and energy tracking. However, evaluations are correlational/benchmarking only (no causal identification), rely on small curated test slices (tinyBenchmarks ~100 examples per task), and results are a cross-sectional snapshot subject to model selection, dataset and measurement choices. Methods Rigormedium — Strengths include participatory value elicitation (advisory board), mixed stakeholders (practitioners + 429 end-users), use of multiple benchmarks (including Dutch-origin or human-verified translations), and reproducible tooling (Code, public overview). Limitations: some benchmarks are small (tinyBenchmarks), the novel HONESTCITYBENCH needs external validation, energy estimates via CodeCarbon can vary with deployment environment, bias measures depend on chosen Dutch-specific datasets and protected characteristics, and RAG / domain-adapted performance appears only mentioned as future/partial work. SampleBenchmarks and evaluations cover more than 30 LLMs (a mix of multilingual and Dutch-specific models such as GEITje, Fietje, GPT-NL and commercial multilingual models); advisory board of nine City of Amsterdam experts (7 respondents to questionnaire); practitioner interviews with 18 City of Amsterdam staff across roles; chatbot user survey with 429 internal users; benchmark datasets include Dutch-created or human-verified translations where possible, Okapi machine-translations, MBBQ/Burema for bias, HONESTCITYBENCH for honesty, tinyBenchmarks (~100 examples per task) for computational tractability, and energy tracking via CodeCarbon. Themesgovernance adoption GeneralizabilityResults are specific to Dutch language models and Dutch governmental context (City of Amsterdam) and may not generalise to other languages or national/local administrations., Model set is a snapshot in time; rapidly evolving LLMs and model updates may change rankings., Benchmark selection (including small tinyBenchmarks samples and newly developed HONESTCITYBENCH) may bias measured performance and may not cover all relevant real-world tasks or RAG deployments., Bias and fairness evaluations are limited to selected protected categories and Dutch-adapted datasets; other harms or demographic groups may be unmeasured., Energy estimates depend on measurement environment and inference setup (API vs self-hosting), limiting direct comparability across deployment modes.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The Grip on LLMs framework identifies six evaluation dimensions for Dutch governmental LLM use: factuality, honesty, social bias, energy consumption, cost, and training data transparency. Governance And Regulation positive Governmental LLM evaluation and model-selection criteria
Reading fidelity high
Study strength medium
n=9
0.18
No single evaluated model performs best across all evaluation dimensions, so governmental model selection requires explicit trade-offs rather than optimization of a single metric. Task Allocation mixed Relative model performance across governmental evaluation dimensions
Reading fidelity high
Study strength medium
more than 30 models
0.18
Higher model quality is consistently associated with greater environmental impact and financial cost. Organizational Efficiency negative Relationship between LLM quality, energy consumption, and deployment cost
Reading fidelity high
Study strength medium
not reported
0.18
Social bias is largely independent of both model quality and the environmental or financial costs associated with models. Ai Safety And Ethics null_result Association between social-bias scores and model quality, energy consumption, and cost
Reading fidelity high
Study strength medium
not reported
0.18
Factuality and honesty are distinct model properties: strong factuality does not necessarily imply strong honesty. Decision Quality mixed Model factual accuracy and willingness to acknowledge uncertainty or inability to answer
Reading fidelity high
Study strength medium
not reported
0.18
Among the seven advisory-board members who completed the value questionnaire, inclusion was the highest-rated value: six members gave it the maximum score of 5. Governance And Regulation positive Priority assigned to inclusion in governmental LLM deployment
Reading fidelity high
Study strength low
n=7
6 out of 7 members assigning it the maximum score of 5
0.09
Factuality was the second-highest-prioritized value among advisory-board respondents, with five of seven members rating it 5 out of 5. Governance And Regulation positive Priority assigned to factuality in governmental LLM deployment
Reading fidelity high
Study strength low
n=7
5 out of 7 members rating factuality 5 out of 5
0.09
Users of the organization’s internal chatbot rated quality and factuality as the most important dimensions, while sustainability and inclusion received substantially lower priority. Worker Satisfaction mixed User-rated priorities for LLM evaluation dimensions
Reading fidelity high
Study strength low
n=429
0.09
Practitioners found existing LLM leaderboards unintuitive and inaccessible, and reported difficulty interpreting raw benchmark scores in relation to specific use cases. Decision Quality negative Interpretability and usability of LLM evaluation information for model selection
Reading fidelity high
Study strength low
n=18
0.09
Advisory-board members considered honesty distinct from factuality because a model that declines to answer appropriately can be preferable to one that produces fluent but unreliable output. Ai Safety And Ethics positive Preference for transparent acknowledgment of model limitations
Reading fidelity high
Study strength low
n=9
0.09

Notes