The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Off‑the‑shelf LLMs remain structurally general-purpose even after customization, and embedding them in government prototypes like Redbox creates persistent governance, procurement and cultural frictions; these dynamics shift public‑sector labor toward oversight and vendor‑management roles and raise regulatory and lock‑in costs that may blunt efficiency gains.

Zero-shot governance
Carlo Perrotta · September 09, 2026 · Journal of Education Policy
openalex descriptive low evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Carlo Perrotta provider ID
A qualitative case study of the UK Redbox prototype argues that the intrinsic general-purpose character of LLMs persists despite domain adaptation, producing enduring technical, political, and cultural challenges that reshape public-sector work, procurement, and governance.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This article interrogates the concept of general-purpose AI governance. General-purpose AI governance – or ‘zero-shot governance’ – refers to a scenario in which instances of domain-agnostic generative AI intervene in policy decisions. This scenario is premised on the potential of the latest generation of AI foundation models to adapt to unfamiliar situations, requiring only relatively agile forms of customisation based on the introduction of domain-specific data. The central argument develops from the infrastructural analysis of Redbox, a now discontinued prototype designed to assist the daily cognitive work of civil servants in the UK. This tool is notable for representing a proof of concept for the integration of off-the-shelf ‘Large Language Models’ (LLMs) in the professional toolkit of policy. This integration is examined through intersecting technical, political, and cultural lenses. A review of the original codebase is followed by the analysis of the political and economic conditions in which Redbox and some its successors emerged. The overarching point emerging from the analysis is that the general-purpose nature of LLMs remains a structural component of the technology that can be mitigated but never ruled out. The conclusion reflects critically on the implications for education policy.

Summary

Main Finding

The article argues that the “general-purpose” character of large language models (LLMs) — their ability to be repurposed across domains with relatively light, domain-specific adaptation — is an enduring structural feature that can be mitigated but not eliminated. Using the UK government prototype Redbox as a case study, the paper shows that integrating off‑the‑shelf LLMs into policy work creates persistent technical, political and cultural dynamics that complicate governance, regulation, and policy practice (including education policy).

Key Points

  • Definition: “General-purpose AI governance” or “zero-shot governance” refers to scenarios where domain-agnostic generative AIs are used to intervene in policy decisions with minimal domain-specific engineering.
  • Case study: Redbox — a discontinued UK prototype for assisting civil servants — serves as a proof of concept for embedding off-the-shelf LLMs into routine policy work.
  • Persistent generality: Even with domain-specific customisation and safeguards, the intrinsic general-purpose capacities of LLMs remain structurally present and can re-emerge in application.
  • Multi-lens analysis: The paper analyses Redbox through technical (codebase and model affordances), political (procurement, vendor relations, institutional incentives), and cultural (professional practices, trust, norms) lenses.
  • Emergence context: The design and deployment of Redbox and successors are shaped by political-economic conditions (e.g., outsourcing, vendor ecosystems, productisation of models), which influence governance outcomes.
  • Education policy: The conclusion issues a critical reflection on how education policy must adapt to the epistemic and civic implications of deploying general-purpose models in public decision-making.

Data & Methods

  • Primary technical analysis: review of the original Redbox codebase and design artifacts to surface architectural choices, customisation layers, and failure modes.
  • Infrastructural/institutional analysis: examination of the political and economic context surrounding Redbox’s emergence — e.g., procurement pathways, vendor relationships, and market incentives.
  • Interdisciplinary framing: synthesis across technical, political, and cultural perspectives to interpret how model generality interacts with institutional practices.
  • Methodological orientation: qualitative, interpretive inquiry grounded in infrastructure studies and socio-technical analysis rather than quantitative evaluation of model performance.

Implications for AI Economics

  • Labour and task allocation: General-purpose models lower the marginal cost of automating a wide range of cognitive tasks, shifting the frontier between human and machine work in public administration and increasing demand for oversight, auditing, and governance skills.
  • Complementarity vs substitution: While LLMs can substitute routine drafting and information synthesis, they create complementary roles (policy stewards, evaluators, prompt engineers) whose remuneration and skill profiles reshape labour markets in government and adjacent sectors.
  • Vendor dynamics & market structure: Off-the-shelf integration encourages reliance on a narrow set of foundation-model providers and middleware vendors, increasing lock‑in risks, bargaining asymmetries, and concentration in the AI supply chain.
  • Regulatory and compliance costs: Mitigating the risks from model generality (alignment checks, audits, domain-specific fine-tuning, provenance tracking) imposes ongoing governance costs that may offset apparent efficiency gains.
  • Public procurement and path dependence: Early prototyping and procurement decisions (like Redbox) can lock institutions into particular architectures and vendor ecosystems, producing long-term economic consequences for competition and innovation trajectories.
  • Externalities and public goods: Misuse, model hallucinations, or poorly governed interventions create negative externalities that justify public investment in oversight, standards, and education to internalise social costs.
  • Education and human capital: There is an economic rationale for targeted upskilling (digital literacy, model governance, critical evaluation) to ensure labour can capture the complementary gains from LLMs and to reduce costs associated with monitoring and correcting automated outputs.

If you want, I can expand any section (e.g., provide a short list of likely governance interventions, or map these implications to specific education-policy prescriptions).

Assessment

Paper Typedescriptive Evidence Strengthlow — The paper is a qualitative, interpretive case study based on a single prototype (Redbox) and documentary/codebase review; it provides rich contextual insights but does not provide quantitative or causal identification of economic effects. Methods Rigormedium — The authors combine direct technical inspection (codebase and design artifacts) with institutional and political-economic analysis, which is appropriate for socio-technical inquiry; however, claims about economic impacts are largely inferential and not supported by systematic empirical measurement or triangulation across multiple comparable cases. SamplePrimary material consists of the Redbox prototype codebase and design artifacts, supplemented by public documents, procurement materials, vendor information, and institutional context analysis; no systematic quantitative dataset or causal inference sample is used. Themesgovernance labor_markets skills_training org_design GeneralizabilitySingle-case study (Redbox) in UK central government limits external validity to other institutions, countries, or private-sector settings., Redbox was a prototype and discontinued — findings may not map to mature, production deployments., Vendor landscape and model capabilities evolve rapidly, so technological specifics may become dated., Interpretive qualitative methods limit ability to generalize effect sizes or causal mechanisms across contexts.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The general-purpose character of large language models is an enduring structural feature that can be mitigated but not eliminated through domain-specific adaptation. Task Allocation mixed Persistence of general-purpose model capabilities under domain-specific customization
Reading fidelity high
Study strength medium
not reported
0.18
Integrating off-the-shelf LLMs into policy work creates persistent technical, political, and cultural dynamics that complicate governance, regulation, and policy practice. Governance And Regulation negative Difficulty of governing and regulating LLM-supported policy work
Reading fidelity high
Study strength medium
not reported
0.18
The Redbox prototype demonstrates the feasibility of embedding off-the-shelf LLMs into routine policy work. Organizational Efficiency positive Feasibility of LLM deployment in routine government policy tasks
Reading fidelity high
Study strength medium
not reported
0.18
The design and deployment of Redbox and its successors are shaped by outsourcing, vendor ecosystems, and the productisation of models, which influence governance outcomes. Governance And Regulation mixed Influence of political-economic conditions on AI governance outcomes
Reading fidelity high
Study strength medium
not reported
0.18
General-purpose models lower the marginal cost of automating a wide range of cognitive tasks in public administration. Task Allocation positive Marginal cost of automating cognitive administrative tasks
Reading fidelity high
Study strength low
not reported
0.09
LLMs may substitute for routine drafting and information synthesis while increasing demand for complementary oversight, auditing, evaluation, and governance roles. Task Allocation mixed Allocation of routine drafting and oversight tasks between humans and LLMs
Reading fidelity high
Study strength speculative
not reported
0.03
Off-the-shelf LLM integration can increase reliance on a narrow set of foundation-model providers and middleware vendors, creating lock-in risks, bargaining asymmetries, and supply-chain concentration. Market Structure negative Concentration and dependence in the AI supply chain
Reading fidelity high
Study strength medium
not reported
0.18
Mitigating risks associated with model generality requires ongoing governance activities, including alignment checks, audits, domain-specific fine-tuning, and provenance tracking, which may offset apparent efficiency gains. Regulatory Compliance negative Governance and compliance costs of deploying general-purpose LLMs
Reading fidelity high
Study strength low
not reported
0.09
Early prototyping and procurement decisions can lock public institutions into particular technical architectures and vendor ecosystems, with long-term consequences for competition and innovation trajectories. Market Structure negative Long-term effects of procurement choices on competition and innovation
Reading fidelity high
Study strength medium
not reported
0.18
The risks of misuse, hallucinations, and poorly governed AI interventions create negative externalities that support public investment in oversight, standards, and education. Ai Safety And Ethics negative Social costs and governance risks from LLM deployment
Reading fidelity high
Study strength speculative
not reported
0.03
Targeted upskilling in digital literacy, model governance, and critical evaluation can help workers capture complementary gains from LLMs and reduce the costs of monitoring and correcting automated outputs. Skill Acquisition positive Worker capability to use, monitor, and critically evaluate LLM outputs
Reading fidelity high
Study strength speculative
not reported
0.03

Notes