0 cumulative citations
View corpus contextGeneral-purpose AI breaks the assumptions of existing policing and public-service governance: its open-ended outputs, accessibility, and persuasive fluency render accuracy, bias measurement, explainability, and accountability frameworks insufficient unless governance explicitly distinguishes GPAI and builds independent safety infrastructure.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextPublic services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.
Summary
Main Finding
The paper argues that the arrival of general-purpose AI (GPAI) — large language models that produce open-ended natural‑language outputs and can be redeployed by prompt alone — fundamentally undermines the safety assumptions and governance tools developed for narrow, purpose‑built AI. In policing (the case study) this mismatch creates a high risk that public‑service AI governance frameworks will fail: accuracy, bias, explainability, and accountability become far harder or impossible to guarantee using existing methods, turning safety from an intrinsic part of system design into an optional, external add‑on. The authors recommend halting operational GPAI deployment in policing until rigorous, independent evidence and a national safety infrastructure exist, preferring technological parsimony and a clear regulatory distinction between narrow AI and GPAI.
Key Points
- Distinct regimes: Narrow (task‑specific) AI enabled tractable safety practices (measurable accuracy, disaggregable bias, feature‑attribution explainability, clear lines of accountability). GPAI breaks those assumptions.
- Four core safety concepts invert under GPAI:
- Accuracy: No fixed standard of correctness for open‑ended outputs; construct validity and measurement fail for many GPAI tasks.
- Bias: Bias becomes a structural property absorbed from huge training corpora and is harder to detect or disaggregate in free‑text judgements.
- Explainability: Fluent rationalizations produce the appearance of explanation without fidelity to model reasoning.
- Accountability: Outputs that convincingly persuade can displace or erode human oversight; the line between tool and decision‑maker blurs.
- Common mitigations in policing (expert evaluation and human‑in‑the‑loop oversight) presuppose narrow‑AI properties and are insufficient when marginal deployment costs fall to prompt writing and local staff adoption.
- “Structural inversion”: GPAI’s accessibility and low marginal deployment cost shift safety assurance from being embedded in system design to being an optional afterthought, increasing risk and reducing oversight feasibility.
- Policy recommendations: explicitly distinguish narrow vs. general‑purpose AI in governance, prefer simple (parsimonious) technology choices, pause GPAI operational use in policing until evidence exists, and build a coordinated, independent national safety infrastructure with authority to evaluate and certify deployments.
- While policing is the detailed case, the authors argue these problems generalize across public services (health, social care, benefits, etc.) where stakes and distributional effects are large.
Data & Methods
- Type: Conceptual and policy analysis / scholarly preprint (not peer‑reviewed yet).
- Unit of analysis: The GPAI model class (rather than any one product or interface) to highlight properties that travel across applications.
- Methods:
- Literature synthesis of AI safety, public‑sector governance guidance, and empirical findings on narrow AI (e.g., recidivism prediction, facial recognition).
- Comparative argument tracing how safety methods developed for narrow AI rely on bounded task definitions and why those methods fail for GPAI.
- Case‑study focus on policing to illustrate practical consequences, referencing legal cases, policy documents (EU AI Act, UK DSIT plans), benchmark studies, and recent model evaluations.
- No original empirical dataset or quantitative experiment; conclusions derive from theoretical analysis + synthesis of existing empirical and policy literature.
Implications for AI Economics
- Misleading cost‑benefit calculations: Standard ROI claims for GPAI in public services (efficiency and labor savings) understate the unmeasured safety, legal, reputational, and remediation costs that arise when accuracy and harms cannot be reliably measured ex ante.
- Negative externalities and market failure: Low marginal cost and broad accessibility create rapid, decentralized adoption (including informal rollouts) that can generate systemic harms not internalized by purchasers — suggesting a case for public intervention, regulation, or insurance mechanisms.
- Procurement and investment strategy:
- Favor technological parsimony: cheaper, narrower systems with measurable performance may produce superior social returns when risk and monitoring costs are included.
- Procurement should require independent evaluation, transparent evidence of task‑specific safety, and mechanisms to internalize liabilities (contractual, insurance, and warranties).
- Public spending priorities:
- The authors’ call for a coordinated national safety infrastructure implies upfront public investment (testing labs, certification bodies, monitoring regimes) that produces public goods: standardized evaluation protocols, reproducible evidence, and regulatory capacity.
- Such investment can reduce information asymmetries and transaction costs for procuring agencies, but it is itself costly and requires clear mandates and independence.
- Labor and substitution effects: While GPAI lowers the cost of automating a wide array of public‑sector tasks, the uncertain quality of outputs and supervision costs may reduce realized labor savings; miscalculated substitution can degrade service quality and increase long‑run expenditures (litigation, corrective interventions).
- Distributional impacts and social welfare: Hard‑to‑detect structural bias in GPAI can produce concentrated harms to vulnerable populations, increasing social costs and potentially requiring targeted mitigation spending. Economic evaluations must account for distributional risks, not just aggregate efficiency gains.
- Regulatory timing matters: A precautionary pause (as recommended for policing) affects adoption curves and market incentives; delaying deployment reallocates private and public investments toward safer, more measurable technologies but may also slow potentially beneficial innovations. Policymakers should weigh dynamic tradeoffs and consider sunset/conditional deployment rules tied to evidence thresholds.
- Role for economic instruments: Taxes, liability rules, procurement weights, and subsidies can be designed to internalize GPAI externalities and steer adoption toward applications where evaluation and accountability are tractable.
Note: This is a preprint by Sam Relins and Daniel Birks (University of Leeds) submitted for publication (not yet peer‑reviewed).
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| General-purpose AI models accept natural-language instructions and can be directed to perform an effectively unbounded range of tasks, unlike narrow AI systems designed for specific tasks. Task Allocation | positive | Range of tasks that an AI model can perform |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The low marginal cost and expertise required to deploy GPAI reduce the barriers to applying AI to new public-service tasks. Adoption Rate | positive | Cost and expertise required to deploy new AI applications |
Reading fidelity
high
Study strength
medium
|
not reported
|
| GPAI models have reported performance approaching that of trained professionals on some legal, medical, scientific, mathematical, and software-engineering benchmarks. Other | positive | Benchmark performance relative to trained professionals |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Nearly 70 countries have adopted national AI strategies. Adoption Rate | positive | National adoption of AI strategies |
Reading fidelity
high
Study strength
medium
|
n=70
Nearly seventy countries
|
| A UK government review estimates that fully digitizing public services could generate more than £45 billion in unrealized annual savings, with AI-driven automation described as a central lever. Organizational Efficiency | positive | Potential annual public-service cost savings |
Reading fidelity
high
Study strength
medium
|
over £45 billion per year
|
| Existing AI safety governance frameworks cannot reliably quantify the accuracy of GPAI outputs across its open-ended task space. Ai Safety And Ethics | negative | Ability to evaluate and certify output accuracy and safety |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Passing mechanically checkable tests does not necessarily establish that GPAI outputs are fit for operational use. Output Quality | negative | Operational fitness of AI-generated outputs after passing automated tests |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Findings about GPAI safety failures generalize poorly across operational contexts because the models have no stable task definition or settled standard of success. Ai Safety And Ethics | negative | Transferability of safety evidence between deployments |
Reading fidelity
high
Study strength
medium
|
not reported
|
| GPAI's open-ended and context-dependent outputs make it difficult to identify a single aggregate test that can reliably detect all relevant harms. Ai Safety And Ethics | negative | Ability of aggregate evaluations to detect safety harms |
Reading fidelity
high
Study strength
low
|
not reported
|
| The paper concludes that existing public-service AI governance frameworks were developed for narrow AI and do not adequately address the safety-assurance problems introduced by GPAI. Governance And Regulation | negative | Adequacy of AI governance frameworks for general-purpose AI |
Reading fidelity
high
Study strength
low
|
not reported
|