3 cumulative citations
View corpus context‘Human in the loop’ mandates often paper over responsibility rather than restore it: legal requirements that a human sign off on automated decisions can shift blame to nominal individuals while concealing institutional and political choices embedded in automation, undermining meaningful oversight.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Human oversight of automated decisions has become regulatory orthodoxy. Scholars and policy makers frequently turn to the human to ameliorate automation’s disruptive potential. But there is little empirical evidence that human oversight or a human in the loop improves decision outcomes. This suggests that the regulatory value of human oversight is more political than it is empirical. This chapter accordingly asks what political work do legal requirements for human oversight perform, and what kinds of political configurations do they enable? Through analysis of cases invoking human in the loop in administrative contexts, this chapter argues that, in the name of recuperating rule of law values, the human in the loop is deployed as a versatile legal technology that enables troubling distributions of accountability and inhibits a richer legal conceptualisation of automated decision systems.
Summary
Main Finding
Jake Goldenfein argues that the “human in the loop” is not merely a technical fix to the governance problems posed by automated decision systems but a highly malleable legal technology. Legal and regulatory invocations of a human in the loop reify unrealistic human/machine binaries and enable politically useful redistributions of accountability — often protecting institutions and officials from substantive responsibility for how automation shapes public decision‑making.
Key Points
- The human in the loop has become a near‑universal regulatory talisman: it promises transparency, dignity, error correction, and a single accountable actor when automated systems make consequential decisions.
- Empirical evidence is mixed or weak that human oversight reliably improves decision quality, reduces bias, or meaningfully increases transparency. Human oversight can be ineffective, uneven, or even facilitate “rubber‑stamping.”
- Two broad responses to that empirical reality are:
- Improve the human side (human‑centred design, training, authority, time, information).
- Reassess the conceptual fit — some scholars argue human and computational epistemic orders are incommensurate, so simple human oversight requirements are conceptually inadequate.
- Goldenfein’s analytic move: instead of only asking whether human oversight works, examine what political and legal work the human in the loop performs. He shows it is used to:
- Mask the political and design choices embedded in automated systems.
- Reassign or obscure accountability away from institutional processes (procurement, design, data curation) onto individual human actors.
- Limit the scope of legal contestation by narrowing what parts of a “decision” count as contestable human acts.
- Case law (e.g., Pintarich v Deputy Commissioner of Taxation) illustrates how courts struggle to define “decision” when processes are partially automated; judicial approaches can either insist on a human mental process or adapt to recognise automated elements — with significant implications for who can be held accountable.
Data & Methods
- Method: qualitative legal and conceptual analysis — close reading of administrative law cases and legal doctrines where human oversight and decision concepts are litigated, combined with synthesis of interdisciplinary literature (legal theory, HCI, STS, public administration).
- Evidence base includes:
- Judicial decisions in administrative law (representative example: Pintarich v Deputy Commissioner of Taxation).
- Regulatory texts and proposals (EU GDPR, EU AI Act proposals, HLEG guidance).
- Empirical and theoretical scholarship on human‑algorithm interaction and oversight (e.g., Green & Chen; De‑Arteaga et al.; Oswald; Raso) and critiques of the human/automation conceptual divide (Amoore, Noll, Andrejevic).
- Analytical emphasis is on tracing how legal doctrines treat the “human” and what governance effects follow, rather than on quantitative evaluation of outcomes.
Implications for AI Economics
- Accountability externalities and moral hazard:
- Treating “human in the loop” as a legal checkbox can create moral hazard: institutions adopt automation while shifting blame to nominal human overseers, reducing incentives to invest in safe system design, data quality, or explainability.
- Economically, this misallocation produces negative externalities (harmful decisions, discrimination, litigation costs) that are not internalised by the designers/contractors who profit from automation.
- Regulatory design and compliance costs:
- Vague human‑oversight mandates generate uncertainty and compliance rents. Firms and procuring agencies can satisfy minimal formal requirements cheaply (e.g., token human review) rather than meaningful oversight, creating regulatory arbitrage opportunities.
- Meaningful oversight is costly — it requires training, time, information flows, and authority. Economists should factor these ongoing costs into cost–benefit estimates for automation in the public sector.
- Labour and task allocation:
- Automation plus nominal human oversight reshapes administrative labour: tasks shift from discretionary judgment to monitoring and rubber‑stamping, changing skill demands, wages, and job quality. This can depress the market value of experienced adjudicative labour while creating precarious oversight roles.
- Market and procurement incentives:
- Procuring agencies may prefer systems that externalise accountability (cheaper deployment, political insulation), skewing market structure toward vendors who optimise for legal/organizational defensibility rather than social welfare or fairness.
- Procurement that fails to allocate liability properly will distort innovation incentives away from safer, more transparent systems.
- Measurement and evaluation needs:
- Economists and policymakers should develop measurable criteria for “meaningful” human oversight (e.g., measurable ability to alter outcomes, time per case, qualifications, information access), and empirically evaluate how oversight affects error rates, bias, and welfare.
- Randomised evaluations and field experiments can help estimate the marginal value of different human oversight designs.
- Policy implications and recommendations:
- Shift regulatory focus from nominal human presence to institutional accountability: hold agencies and vendors responsible for design, data, auditability, and outcomes through liability rules, mandatory audits, and disclosure regimes.
- Align incentives: use liability, procurement terms, and performance‑based contracting to internalise costs of harms and reward investments in explainability and robustness.
- Cost out meaningful oversight: when modelling automation adoption, include realistic recurrent costs for training, time, and information systems needed for genuine human intervention.
- Monitor labour market impacts: anticipate task reallocation and design retraining/compensation policies for affected administrative workers.
- Use clear, operational definitions of what constitutes a “decision” and what counts as reversible/contestable elements, so legal and economic analyses can converge on measurable intervention points.
Short takeaway: For economists, Goldenfein’s chapter reframes the human in the loop not as a neutral technical parameter but as a legal and political instrument that shapes incentives, distributes costs and liabilities, and therefore materially influences the economic calculus of automation in the public sector. Policy and economic models should therefore replace checkbox notions of oversight with empirically anchored, incentive‑aware institutional rules.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The human-in-the-loop has become a central regulatory approach intended to preserve human values and accountability in automated decision-making. Governance And Regulation | positive | Preservation of human values and accountability in automated decisions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The human-in-the-loop paradigm can enable problematic distributions of accountability and inhibit richer legal conceptualisations of automated decision systems. Governance And Regulation | negative | Distribution and conceptualisation of legal accountability for automated decisions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Human oversight of data-driven judicial decisions does little to mitigate discrimination. Inequality | negative | Discrimination in data-driven judicial decisions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Human intervention against malfunctioning algorithmic risk-prediction tools in child-welfare decisions has limited effectiveness, with results that are mixed. Decision Quality | mixed | Human capacity to detect or correct erroneous algorithmic risk scores |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Across the cited research, the effect of human-in-the-loop oversight on decision quality is uneven. Decision Quality | mixed | Decision quality under human oversight of automated systems |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Human-in-the-loop requirements can assign responsibility to individual decision-makers even when those individuals lack control over the outcomes produced by automated systems. Governance And Regulation | negative | Allocation of responsibility and accountability for automated decisions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In Pintarich v Deputy Commissioner of Taxation, the Australian Federal Court majority held that a legally valid decision requires both a mental process of reaching a conclusion and an objective manifestation of that conclusion. Governance And Regulation | positive | Legal recognition of an administrative decision |
Reading fidelity
high
Study strength
high
|
not reported
|
| In Pintarich, the majority found that an automated letter was insufficient to constitute a legal decision because it was not accompanied by the required human cognitive process. Governance And Regulation | negative | Legal validity and reviewability of an automated administrative decision |
Reading fidelity
high
Study strength
high
|
not reported
|