0 cumulative citations
View corpus contextAmsterdam’s evaluation finds no one-size-fits-all LLM for government: models that are more accurate tend to be costlier and more energy-intensive, while bias levels vary independently, meaning municipal procurement must balance quality, sustainability and fairness.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.
Summary
Main Finding
No single LLM is optimal for Dutch governmental use — measurable trade-offs are unavoidable. Higher-quality models (factuality/honesty) tend to incur greater energy use and financial cost, while social bias behaves largely independently of those trade-offs. Factuality and honesty are distinct properties: a model can answer correctly but still fail to acknowledge its limits.
Key Points
- Framework developed with City of Amsterdam stakeholders (advisory board, practitioner interviews, and a 429-user chatbot survey) to translate organisational values into measurable evaluation dimensions.
- Six evaluation dimensions prioritised for governmental deployment: factuality, honesty, social bias, energy consumption (sustainability), cost, and training-data transparency. Two practical use cases were emphasised: text simplification and summarisation.
- Large-scale, Dutch-focused benchmark suite applied to 30+ multilingual and Dutch-specific models (including GEITje, Fietje, GPT‑NL and others).
- Empirical findings:
- No model dominates across all dimensions.
- Quality (factuality/honesty) correlates positively with energy consumption and cost.
- Social bias does not correlate strongly with either cost or energy; it requires separate measurement and mitigation.
- High factuality does not imply high honesty — they must be measured separately.
- Usability: results are translated into a user-friendly, non-technical leaderboard and public overview to support procurement and policy decisions.
- Code and interactive overview released publicly (repository and web dashboard).
Data & Methods
- Stakeholder input:
- Advisory board: 9 invited experts (7 responded to the questionnaire). Process: open questionnaire → deep-dive sessions → iterative feedback.
- Practitioner interviews: 18 staff across roles (data scientists, product owners, managers, policymakers).
- End-user survey: 429 internal chatbot users (prioritised factuality/quality more highly than advisory board).
- Benchmarks and operationalisation:
- Factuality: Dutch or verified human-translated knowledge benchmarks; evaluations economised via tinyBenchmarks (≈100 curated examples per benchmark) to scale across many models.
- Honesty: measured with HONESTCITYBENCH (developed alongside this project) to capture whether models acknowledge limitations/refuse when appropriate.
- Social bias: used Dutch-adapted bias datasets (e.g., MBBQ, Burema) covering protected characteristics (age, origin, disability, gender).
- Sustainability (energy): inference energy tracked with CodeCarbon to estimate environmental impact.
- Cost: estimated from API pricing and self-hosting operational costs to reflect deployment budget impacts.
- Training-data transparency: categorical/qualitative assessment of model license and available documentation on data provenance and collection practices.
- Model coverage: over 30 Dutch-specific and multilingual generative models were evaluated, enabling cross-model comparisons.
- Presentation: technical scores converted into interpretable categorical ratings (1–5 scales, coloured pills, and normalised bar charts) for non-technical stakeholders.
Implications for AI Economics
- Multi-dimensional procurement decisions: Governments should avoid single-metric selection (e.g., highest accuracy). Procurement frameworks must operationalise explicit trade-offs (quality vs cost vs carbon).
- Internalisation of externalities: Energy (carbon) costs correlate with higher-performing models — public buyers should account for environmental externalities in total-cost-of-ownership and may require carbon budget constraints or prefer energy-efficient alternatives for low-risk tasks.
- Pricing and market structure:
- Higher-quality models commanding higher inference costs could reinforce market concentration (providers able to fund large-scale models). Public procurement that values openness and training-data transparency may help sustain competition and local/national model development.
- There is room for differentiated model choice by use case: lower-cost, lower-energy models may suffice for routine, non-critical tasks (e.g., basic drafting, summarisation), while high-criticality tasks require higher-quality models plus stronger oversight.
- Bias requires separate regulation and monitoring: since social bias is not strongly tied to cost or energy, cost-based procurement alone won’t mitigate fairness risks. Regular bias audits and remediation investments are necessary.
- Honesty and trustworthiness as economic factors: models that better acknowledge uncertainties reduce downstream verification costs and potential liability; quantifying honesty can inform expected verification overheads and public-trust trade-offs.
- Support for local-language and open models: investing in Dutch-specific and transparent models can reduce dependency on expensive multilingual commercial offerings, potentially lowering both monetary and environmental costs while improving data-governance alignment (GDPR, copyright).
- Policy recommendations for governments:
- Adopt multi-criteria decision frameworks that explicitly weight factuality, honesty, bias, cost, energy, and data transparency.
- Require reporting of inference energy and training-data provenance in procurement bids.
- Fund or procure energy-efficient models and tools for retrieval-augmented systems (which may reduce reliance on large parametric models).
- Mandate independent bias and honesty evaluations as part of deployment approvals.
If you want, I can (a) extract the concrete numeric results for specific models from the paper and present them in a compact table, or (b) draft a short procurement checklist based on the framework for municipal buyers. Which would be most useful?
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The Grip on LLMs framework identifies six evaluation dimensions for Dutch governmental LLM use: factuality, honesty, social bias, energy consumption, cost, and training data transparency. Governance And Regulation | positive | Governmental LLM evaluation and model-selection criteria |
Reading fidelity
high
Study strength
medium
|
n=9
|
| No single evaluated model performs best across all evaluation dimensions, so governmental model selection requires explicit trade-offs rather than optimization of a single metric. Task Allocation | mixed | Relative model performance across governmental evaluation dimensions |
Reading fidelity
high
Study strength
medium
|
more than 30 models
|
| Higher model quality is consistently associated with greater environmental impact and financial cost. Organizational Efficiency | negative | Relationship between LLM quality, energy consumption, and deployment cost |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Social bias is largely independent of both model quality and the environmental or financial costs associated with models. Ai Safety And Ethics | null_result | Association between social-bias scores and model quality, energy consumption, and cost |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Factuality and honesty are distinct model properties: strong factuality does not necessarily imply strong honesty. Decision Quality | mixed | Model factual accuracy and willingness to acknowledge uncertainty or inability to answer |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Among the seven advisory-board members who completed the value questionnaire, inclusion was the highest-rated value: six members gave it the maximum score of 5. Governance And Regulation | positive | Priority assigned to inclusion in governmental LLM deployment |
Reading fidelity
high
Study strength
low
|
n=7
6 out of 7 members assigning it the maximum score of 5
|
| Factuality was the second-highest-prioritized value among advisory-board respondents, with five of seven members rating it 5 out of 5. Governance And Regulation | positive | Priority assigned to factuality in governmental LLM deployment |
Reading fidelity
high
Study strength
low
|
n=7
5 out of 7 members rating factuality 5 out of 5
|
| Users of the organization’s internal chatbot rated quality and factuality as the most important dimensions, while sustainability and inclusion received substantially lower priority. Worker Satisfaction | mixed | User-rated priorities for LLM evaluation dimensions |
Reading fidelity
high
Study strength
low
|
n=429
|
| Practitioners found existing LLM leaderboards unintuitive and inaccessible, and reported difficulty interpreting raw benchmark scores in relation to specific use cases. Decision Quality | negative | Interpretability and usability of LLM evaluation information for model selection |
Reading fidelity
high
Study strength
low
|
n=18
|
| Advisory-board members considered honesty distinct from factuality because a model that declines to answer appropriately can be preferable to one that produces fluent but unreliable output. Ai Safety And Ethics | positive | Preference for transparent acknowledgment of model limitations |
Reading fidelity
high
Study strength
low
|
n=9
|