The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

General-purpose AI breaks the assumptions of existing policing and public-service governance: its open-ended outputs, accessibility, and persuasive fluency render accuracy, bias measurement, explainability, and accountability frameworks insufficient unless governance explicitly distinguishes GPAI and builds independent safety infrastructure.

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Relins, Sam, Birks, Daniel · July 28, 2026 · arXiv (Cornell University)
openalex theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Relins, Sam provider ID
  2. Birks, Daniel provider ID

Semantic Scholar

Latest observation:

  1. Samuel Relins provider ID
  2. Dan Birks provider ID
The paper argues that general-purpose AI's generality, accessibility, and low deployment cost fundamentally undermine the measurement- and task-based safety assurances that underpin current public-service (especially policing) AI governance, and recommends distinct regulatory treatment, technological parsimony, a deployment pause for GPAI in policing, and a national independent safety infrastructure.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.

Summary

Main Finding

The paper argues that the arrival of general-purpose AI (GPAI) — large language models that produce open-ended natural‑language outputs and can be redeployed by prompt alone — fundamentally undermines the safety assumptions and governance tools developed for narrow, purpose‑built AI. In policing (the case study) this mismatch creates a high risk that public‑service AI governance frameworks will fail: accuracy, bias, explainability, and accountability become far harder or impossible to guarantee using existing methods, turning safety from an intrinsic part of system design into an optional, external add‑on. The authors recommend halting operational GPAI deployment in policing until rigorous, independent evidence and a national safety infrastructure exist, preferring technological parsimony and a clear regulatory distinction between narrow AI and GPAI.

Key Points

  • Distinct regimes: Narrow (task‑specific) AI enabled tractable safety practices (measurable accuracy, disaggregable bias, feature‑attribution explainability, clear lines of accountability). GPAI breaks those assumptions.
  • Four core safety concepts invert under GPAI:
    • Accuracy: No fixed standard of correctness for open‑ended outputs; construct validity and measurement fail for many GPAI tasks.
    • Bias: Bias becomes a structural property absorbed from huge training corpora and is harder to detect or disaggregate in free‑text judgements.
    • Explainability: Fluent rationalizations produce the appearance of explanation without fidelity to model reasoning.
    • Accountability: Outputs that convincingly persuade can displace or erode human oversight; the line between tool and decision‑maker blurs.
  • Common mitigations in policing (expert evaluation and human‑in‑the‑loop oversight) presuppose narrow‑AI properties and are insufficient when marginal deployment costs fall to prompt writing and local staff adoption.
  • “Structural inversion”: GPAI’s accessibility and low marginal deployment cost shift safety assurance from being embedded in system design to being an optional afterthought, increasing risk and reducing oversight feasibility.
  • Policy recommendations: explicitly distinguish narrow vs. general‑purpose AI in governance, prefer simple (parsimonious) technology choices, pause GPAI operational use in policing until evidence exists, and build a coordinated, independent national safety infrastructure with authority to evaluate and certify deployments.
  • While policing is the detailed case, the authors argue these problems generalize across public services (health, social care, benefits, etc.) where stakes and distributional effects are large.

Data & Methods

  • Type: Conceptual and policy analysis / scholarly preprint (not peer‑reviewed yet).
  • Unit of analysis: The GPAI model class (rather than any one product or interface) to highlight properties that travel across applications.
  • Methods:
    • Literature synthesis of AI safety, public‑sector governance guidance, and empirical findings on narrow AI (e.g., recidivism prediction, facial recognition).
    • Comparative argument tracing how safety methods developed for narrow AI rely on bounded task definitions and why those methods fail for GPAI.
    • Case‑study focus on policing to illustrate practical consequences, referencing legal cases, policy documents (EU AI Act, UK DSIT plans), benchmark studies, and recent model evaluations.
  • No original empirical dataset or quantitative experiment; conclusions derive from theoretical analysis + synthesis of existing empirical and policy literature.

Implications for AI Economics

  • Misleading cost‑benefit calculations: Standard ROI claims for GPAI in public services (efficiency and labor savings) understate the unmeasured safety, legal, reputational, and remediation costs that arise when accuracy and harms cannot be reliably measured ex ante.
  • Negative externalities and market failure: Low marginal cost and broad accessibility create rapid, decentralized adoption (including informal rollouts) that can generate systemic harms not internalized by purchasers — suggesting a case for public intervention, regulation, or insurance mechanisms.
  • Procurement and investment strategy:
    • Favor technological parsimony: cheaper, narrower systems with measurable performance may produce superior social returns when risk and monitoring costs are included.
    • Procurement should require independent evaluation, transparent evidence of task‑specific safety, and mechanisms to internalize liabilities (contractual, insurance, and warranties).
  • Public spending priorities:
    • The authors’ call for a coordinated national safety infrastructure implies upfront public investment (testing labs, certification bodies, monitoring regimes) that produces public goods: standardized evaluation protocols, reproducible evidence, and regulatory capacity.
    • Such investment can reduce information asymmetries and transaction costs for procuring agencies, but it is itself costly and requires clear mandates and independence.
  • Labor and substitution effects: While GPAI lowers the cost of automating a wide array of public‑sector tasks, the uncertain quality of outputs and supervision costs may reduce realized labor savings; miscalculated substitution can degrade service quality and increase long‑run expenditures (litigation, corrective interventions).
  • Distributional impacts and social welfare: Hard‑to‑detect structural bias in GPAI can produce concentrated harms to vulnerable populations, increasing social costs and potentially requiring targeted mitigation spending. Economic evaluations must account for distributional risks, not just aggregate efficiency gains.
  • Regulatory timing matters: A precautionary pause (as recommended for policing) affects adoption curves and market incentives; delaying deployment reallocates private and public investments toward safer, more measurable technologies but may also slow potentially beneficial innovations. Policymakers should weigh dynamic tradeoffs and consider sunset/conditional deployment rules tied to evidence thresholds.
  • Role for economic instruments: Taxes, liability rules, procurement weights, and subsidies can be designed to internalize GPAI externalities and steer adoption toward applications where evaluation and accountability are tractable.

Note: This is a preprint by Sam Relins and Daniel Birks (University of Leeds) submitted for publication (not yet peer‑reviewed).

Assessment

Paper Typetheoretical Evidence Strengthn/a — Paper is a conceptual, policy-argument preprint that synthesizes prior literature and regulatory documents rather than reporting original empirical causal evidence or quantitative analysis. Methods Rigorn/a — The manuscript advances a reasoned theoretical argument grounded in existing literature and policy sources rather than applying an empirical identification strategy or quantitative methods; the argumentation appears systematic but is not subject to empirical validation within the paper. SampleNo original empirical sample or dataset; the paper is a conceptual analysis drawing on prior academic literature, policy documents, government reviews, and policing case studies (primarily UK-focused references). Themesgovernance adoption productivity org_design human_ai_collab GeneralizabilityArgument is largely conceptual and illustrative using policing; implications for other public services are plausible but not empirically demonstrated., UK- and policing-centric sources may limit direct transferability to other national legal/regulatory contexts and different public services., Claims depend on the current technical characteristics of GPAI and may change as models, toolchains, or evaluation methods evolve., Absence of empirical testing means operational impacts and magnitudes are speculative rather than measured.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
General-purpose AI models accept natural-language instructions and can be directed to perform an effectively unbounded range of tasks, unlike narrow AI systems designed for specific tasks. Task Allocation positive Range of tasks that an AI model can perform
Reading fidelity high
Study strength medium
not reported
0.12
The low marginal cost and expertise required to deploy GPAI reduce the barriers to applying AI to new public-service tasks. Adoption Rate positive Cost and expertise required to deploy new AI applications
Reading fidelity high
Study strength medium
not reported
0.12
GPAI models have reported performance approaching that of trained professionals on some legal, medical, scientific, mathematical, and software-engineering benchmarks. Other positive Benchmark performance relative to trained professionals
Reading fidelity high
Study strength medium
not reported
0.12
Nearly 70 countries have adopted national AI strategies. Adoption Rate positive National adoption of AI strategies
Reading fidelity high
Study strength medium
n=70
Nearly seventy countries
0.12
A UK government review estimates that fully digitizing public services could generate more than £45 billion in unrealized annual savings, with AI-driven automation described as a central lever. Organizational Efficiency positive Potential annual public-service cost savings
Reading fidelity high
Study strength medium
over £45 billion per year
0.12
Existing AI safety governance frameworks cannot reliably quantify the accuracy of GPAI outputs across its open-ended task space. Ai Safety And Ethics negative Ability to evaluate and certify output accuracy and safety
Reading fidelity high
Study strength medium
not reported
0.12
Passing mechanically checkable tests does not necessarily establish that GPAI outputs are fit for operational use. Output Quality negative Operational fitness of AI-generated outputs after passing automated tests
Reading fidelity high
Study strength medium
not reported
0.12
Findings about GPAI safety failures generalize poorly across operational contexts because the models have no stable task definition or settled standard of success. Ai Safety And Ethics negative Transferability of safety evidence between deployments
Reading fidelity high
Study strength medium
not reported
0.12
GPAI's open-ended and context-dependent outputs make it difficult to identify a single aggregate test that can reliably detect all relevant harms. Ai Safety And Ethics negative Ability of aggregate evaluations to detect safety harms
Reading fidelity high
Study strength low
not reported
0.06
The paper concludes that existing public-service AI governance frameworks were developed for narrow AI and do not adequately address the safety-assurance problems introduced by GPAI. Governance And Regulation negative Adequacy of AI governance frameworks for general-purpose AI
Reading fidelity high
Study strength low
not reported
0.06

Notes