0 cumulative citations
View corpus contextAccountability for large language models remains fragmented: a systematic review maps technical, governance and documentation mechanisms onto the EU AI Act, NIST and ISO standards and proposes a Layered Accountability Architecture (LAAF), but finds human oversight, shared metrics and empirical validation are persistently under-specified.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility be traced, explained, and acted upon? Following PRISMA guidance, five databases were searched from January 2022 to March 2026 against four review questions; of 4,512 records identified, 122 primary studies were included, together with 12 regulatory and standards documents analysed as primary sources. The review consolidates a sociotechnical account of accountability as an actor-forum relation resolved into five dimensions, and synthesises mechanisms across four families: technical controls, human oversight, organisational governance, and documentation and traceability, each with a maturity assessment. The corpus is read through a four-layer classification device spanning provenance, application logic, human oversight, and governance and redress, cross-cut by traceability, role clarity, and continuous monitoring. Both are mapped onto the EU AI Act, whose high-risk obligations have applied since 2 August 2026, the NIST AI RMF with its Generative AI Profile, ISO/IEC 42001, and sectoral guidance in healthcare, consumer finance, education, and the public sector. Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation, alongside five structural tensions that no surveyed instrument resolves. The review closes by consolidating the classification device into an integrated accountability architecture, LAAF, with cybersecurity aligned to the OWASP LLM Top 10 (2025); it is a synthesis of the surveyed evidence rather than a validated artefact.
Summary
Main Finding
The paper performs a PRISMA-style systematic review (Jan 2022–Mar 2026) of accountability for LLM-based applications and synthesises the evidence into a layered accountability architecture framework (LAAF). LAAF organises accountability across four operational layers (provenance, application logic, human oversight, governance & redress), cross-cut by traceability, role-clarity, and continuous monitoring, and maps mechanisms (technical controls, human oversight, organisational governance, documentation & traceability) onto binding instruments (EU AI Act, NIST AI RMF GenAI Profile, ISO/IEC 42001) and sectoral guidance. The review finds that while many technical mechanisms exist, major gaps persist—particularly under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation—and offers LAAF as an integrated, evidence-based architecture (a synthesis, not yet validated).
Key Points
-
Scope and corpus
- Searched five databases for Jan 2022–Mar 2026; identified 4,512 records.
- Included 122 primary studies and 12 regulatory/standards documents as primary sources.
- Four review questions: conceptualisation of accountability; mechanisms and maturity; mapping onto regulatory landscape; synthesis of patterns/gaps into an architecture.
-
Conceptual contributions
- Defines accountability as an actor–forum relation resolvable into five analytical dimensions (paper details these dimensions and failure modes).
- Distinguishes accountability from transparency, explainability, and liability—each necessary but insufficient.
-
Layered classification (LAAF core)
- Four layers: 1) Provenance (data, training lineage), 2) Application logic (retrieval, prompts, system prompts, grounding), 3) Human oversight (who reviews, authority, procedural rules), 4) Governance & redress (organisational policy, incident handling, legal/regulatory interfaces).
- Three cross-cutting properties: traceability (audit trails, evidence), role clarity (named accountable actors across provider/integrator/deployer/overseer), continuous monitoring (runtime detection, change-control).
-
Mechanism families and maturity
- Technical controls: detection (sample dispersion, alignment with references), grounding (retrieval-augmented generation), guardrails, automated verification — generally the most mature.
- Human oversight: frequently invoked but under-specified (which humans, what evidence, authority to overrule) — low maturity.
- Organisational governance: risk management, change control, incident response — emerging maturity with regulatory anchors.
- Documentation & traceability: provenance records, logging, model cards/audit trails — moderate maturity but inconsistent standards.
- Each family is assessed against maturity and regulatory anchors (EU AI Act, NIST, ISO).
-
Regulatory mapping
- Mapped literature/mechanisms onto EU AI Act (high-risk obligations in force since 2 Aug 2026), NIST AI RMF GenAI Profile, ISO/IEC 42001, and sectoral guidance (healthcare, consumer finance, education, public sector).
- Identified areas where existing literature supports regulatory requirements and where gaps remain.
-
Persistent gaps and tensions
- Four persistent gaps: under-specified human oversight; absence of shared accountability metrics; disciplinary disconnection between technical and governance literatures; limited empirical evaluation of mechanisms.
- Five structural tensions (not fully resolved by surveyed instruments), e.g., between model updates (dynamic systems) and the need for stable auditability; between provider vs deployer responsibilities; between transparency and IP/privacy; between standardized metrics vs domain-specific needs; between automated detection and human adjudication.
-
Hallucinations and formal constraints
- Argues hallucination is systemic to probabilistic LLM generation: formal lower bounds (e.g., single-occurrence facts) and computability results imply hallucination cannot be eliminated—thus must be governed.
- Detection formalisms discussed: dispersion/entropy across multiple samples, alignment score to external references with calibrated thresholds (τ), and combinations; threshold calibration is an accountability decision (trade-off between false positives/negatives).
-
Output
- Proposes LAAF as an integrated accountability architecture, aligned to cybersecurity concerns (OWASP LLM Top 10), synthesising evidence rather than presenting a validated artefact.
Data & Methods
-
Review design
- PRISMA-inspired systematic review covering Jan 2022–Mar 2026.
- Search across five academic/technical databases (authors’ exact list not in this summary but the paper documents them).
- Inclusion: studies addressing accountability mechanisms, conceptual treatments, technical detection/mitigation, governance, and regulatory instruments relevant to LLM deployments.
- Result: 4,512 identified records → 122 primary studies included + 12 regulatory/standards documents analysed as primary sources.
-
Analytical approach
- Conceptual synthesis: consolidated accountability into five dimensions and mapped failure modes (Section 5 of paper).
- Mechanism synthesis: grouped mechanisms into four families and assessed maturity, regulatory anchors, and sectoral applicability (Section 6).
- Regulatory mapping: comparative mapping to EU AI Act, NIST AI RMF (GenAI Profile), ISO/IEC 42001, and sectoral guidance (Section 7).
- Identification of persistent gaps and tensions via cross-corpus analysis (Section 8).
- Consolidation into LAAF: integrated the above into layered architecture (Section 9); explicitly stated as a synthesis (not validated empirically).
-
Formal/theoretical support
- Incorporated formal results about inevitability of some hallucinations (references in paper) and probabilistic detection formalisms (entropy/dispersion, alignment scores, threshold calibration).
Implications for AI Economics
-
Liability, risk allocation, and incentives
- Clear role-clarity requirements (provider vs integrator vs deployer) will shift where economic liability and compliance costs sit. Firms may re-arrange contracts, warranties, and terms of service to allocate legal and financial risk.
- Under-specified human oversight creates regulatory and contractual uncertainty; economic actors will internalise this through higher premiums, stricter contracting, or avoidance of high-risk deployments.
-
Compliance costs and barriers to entry
- Documentation, continuous monitoring, provenance logging, and certification to meet EU AI Act and similar standards impose fixed and ongoing costs. These create economies of scale favoring larger incumbents and can raise barriers to market entry for smaller firms/startups.
- Conversely, specialised compliance-as-a-service providers and certification markets may emerge, creating new economic niches.
-
Insurance and risk-pooling markets
- Absence of shared accountability metrics complicates insurance underwriting for LLM-caused harms. Development of standard metrics (e.g., calibrated hallucination thresholds, demonstrable oversight processes) would reduce information asymmetry and enable scalable insurance products.
- Until such metrics exist, insurance costs will be high or cover limited perils; insurers may require demonstrable adherence to architectures like LAAF, certified monitoring, and documented human-in-loop procedures.
-
Product design and competitive strategy
- Providers may differentiate by offering stronger provenance, explainability, and certified alignment tooling—creating product-market segmentation (safety/compliance premium).
- Deployers (e.g., hospitals, banks) will trade off model capability vs. interpretability and auditability; high-regulation sectors will prefer integrations that maximise traceability even at some performance cost.
-
Dynamic update externalities and platform effects
- The dynamic nature of LLMs (updates to models, retrieval corpora, prompts) produces externalities: a provider update may change downstream deployer risk. Contracts and platform governance will need mechanisms for notification, rollback, and coordinated change-control—economic frictions that can slow innovation or induce conservative update policies.
- Large platform providers may internalise these coordination costs better than small vendors, reinforcing winner-take-most dynamics.
-
Measurement, metrics, and market signalling
- Calibrated thresholds and measurable alignment scores become economic signals of trustworthiness. Firms able to credibly report standardised metrics can command higher prices or lower capital/insurance costs.
- The current lack of shared metrics is a market failure: private actors have weak incentives to invest early in costly, non-excludable metrics. Regulatory intervention or industry consortia may be necessary to create public goods (standard metrics, benchmark datasets).
-
Sectoral heterogeneity and social welfare
- Sector-specific guidance (healthcare, finance, education, public sector) implies differential compliance burdens and welfare implications. In health and finance, the social cost of errors is high—leading to stricter oversight and higher costs that may reduce adoption speed but increase safety. In less regulated sectors, faster deployment may yield consumer surplus but higher systemic risk.
- Policymakers must weigh the trade-off between innovation diffusion and the social cost of harms when calibrating sectoral regulatory strictness.
-
Research & policy priorities for economics
- Prioritise development of shared accountability metrics and standard benchmarking that are interoperable across sectors—this reduces transaction costs and insurance friction.
- Study how compliance costs affect market structure, especially entry and concentration, and design policy measures (e.g., subsidies, standardisation support) to mitigate anti-competitive effects.
- Explore contract designs and liability rules (e.g., mandatory disclosure, strict liability for deployers in certain contexts) that align incentives across providers, integrators, and deployers.
- Model the equilibrium effects of mandatory provenance/traceability requirements on innovation rates, platform dominance, and consumer welfare.
-
Practical takeaways for economic actors
- Firms deploying LLMs should anticipate and budget for ongoing monitoring, provenance logging, and human-in-loop procedures; these are not one-time compliance tasks.
- Early investment in traceability and standardised metrics can yield lower insurance and regulatory friction costs later, and can be a credible market signal.
- Smaller actors should consider partnering with compliance-specialist providers or using certified platforms to manage the fixed costs of compliance.
Note: LAAF is presented as a synthesis of surveyed evidence rather than a validated operational artefact. The paper emphasises the need for empirical evaluation and shared metrics—both essential for translating the architecture into workable economic incentives and market mechanisms.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The review identified 4,512 records and included 122 primary studies, along with 12 regulatory and standards documents analyzed as primary sources. Governance And Regulation | positive | Size and composition of the reviewed evidence corpus |
Reading fidelity
high
Study strength
medium
|
n=4512
|
| The reviewed literature has four persistent accountability gaps: under-specified human oversight, a lack of shared accountability metrics, disconnection between technical and governance disciplines, and limited empirical evaluation of accountability mechanisms. Governance And Regulation | negative | Completeness and empirical maturity of accountability practices |
Reading fidelity
high
Study strength
medium
|
n=122
|
| The review concludes that accountability for LLM applications should be treated as both a lifecycle property and a distributed property allocated across named actors. Governance And Regulation | positive | Clarity and continuity of responsibility assignment across LLM deployment |
Reading fidelity
high
Study strength
medium
|
n=122
|
| The surveyed literature does not support a defensible pooled estimate of hallucination frequency because the studies use incompatible definitions, prompts, and domains. Error Rate | null_result | Pooled hallucination frequency |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The review characterizes hallucination as a systemic property of probabilistic generation rather than merely a transient defect of particular model releases. Error Rate | negative | Reliability and factuality of LLM outputs |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Hallucination-detection approaches based on response dispersion can provide sample-level evidence of hallucination without access to model internals, but they cannot detect confident, consistent errors. Error Rate | mixed | Hallucination detection capability |
Reading fidelity
high
Study strength
medium
|
n=10
|
| The proposed Layered Accountability Architecture Framework (LAAF) is a synthesis of the reviewed evidence and has not been validated as an artefact. Governance And Regulation | positive | Integrated accountability architecture for LLM applications |
Reading fidelity
high
Study strength
speculative
|
n=122
|