The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Accountability for large language models remains fragmented: a systematic review maps technical, governance and documentation mechanisms onto the EU AI Act, NIST and ISO standards and proposes a Layered Accountability Architecture (LAAF), but finds human oversight, shared metrics and empirical validation are persistently under-specified.

LAAF: A Layered Accountability Architecture Framework for LLM Applications
Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi, Muhammad Waqas, George Loukas, Alireza Jolfaei, Lucas Cordeiro, Pierre Dantas · August 27, 2026
arxiv review_meta medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Prachi Chaturvedi unresolved corpus identity
  2. Shahnawaz Ahmad unresolved corpus identity
  3. Ehsan Nowroozi unresolved corpus identity
  4. Muhammad Waqas unresolved corpus identity
  5. George Loukas unresolved corpus identity
  6. Alireza Jolfaei unresolved corpus identity
  7. Lucas Cordeiro unresolved corpus identity
  8. Pierre Dantas unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Prachi Chaturvedi provider ID
  2. Shahnawaz Ahmad unresolved corpus identity
  3. Ehsan Nowroozi provider ID
  4. Muhammad Waqas provider ID
  5. G. Loukas provider ID
  6. A. Jolfaei provider ID
  7. Lucas C. Cordeiro provider ID
  8. Pierre V. Dantas provider ID
This systematic review synthesises conceptual and technical accountability mechanisms for LLM deployments, maps them onto major regulatory instruments (EU AI Act, NIST, ISO), identifies persistent gaps (under-specified human oversight, absent shared metrics, disciplinary disconnection, limited empirical evaluation), and proposes the Layered Accountability Architecture Framework (LAAF) as an integrative synthesis.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility be traced, explained, and acted upon? Following PRISMA guidance, five databases were searched from January 2022 to March 2026 against four review questions; of 4,512 records identified, 122 primary studies were included, together with 12 regulatory and standards documents analysed as primary sources. The review consolidates a sociotechnical account of accountability as an actor-forum relation resolved into five dimensions, and synthesises mechanisms across four families: technical controls, human oversight, organisational governance, and documentation and traceability, each with a maturity assessment. The corpus is read through a four-layer classification device spanning provenance, application logic, human oversight, and governance and redress, cross-cut by traceability, role clarity, and continuous monitoring. Both are mapped onto the EU AI Act, whose high-risk obligations have applied since 2 August 2026, the NIST AI RMF with its Generative AI Profile, ISO/IEC 42001, and sectoral guidance in healthcare, consumer finance, education, and the public sector. Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation, alongside five structural tensions that no surveyed instrument resolves. The review closes by consolidating the classification device into an integrated accountability architecture, LAAF, with cybersecurity aligned to the OWASP LLM Top 10 (2025); it is a synthesis of the surveyed evidence rather than a validated artefact.

Summary

Main Finding

The paper performs a PRISMA-style systematic review (Jan 2022–Mar 2026) of accountability for LLM-based applications and synthesises the evidence into a layered accountability architecture framework (LAAF). LAAF organises accountability across four operational layers (provenance, application logic, human oversight, governance & redress), cross-cut by traceability, role-clarity, and continuous monitoring, and maps mechanisms (technical controls, human oversight, organisational governance, documentation & traceability) onto binding instruments (EU AI Act, NIST AI RMF GenAI Profile, ISO/IEC 42001) and sectoral guidance. The review finds that while many technical mechanisms exist, major gaps persist—particularly under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation—and offers LAAF as an integrated, evidence-based architecture (a synthesis, not yet validated).

Key Points

  • Scope and corpus

    • Searched five databases for Jan 2022–Mar 2026; identified 4,512 records.
    • Included 122 primary studies and 12 regulatory/standards documents as primary sources.
    • Four review questions: conceptualisation of accountability; mechanisms and maturity; mapping onto regulatory landscape; synthesis of patterns/gaps into an architecture.
  • Conceptual contributions

    • Defines accountability as an actor–forum relation resolvable into five analytical dimensions (paper details these dimensions and failure modes).
    • Distinguishes accountability from transparency, explainability, and liability—each necessary but insufficient.
  • Layered classification (LAAF core)

    • Four layers: 1) Provenance (data, training lineage), 2) Application logic (retrieval, prompts, system prompts, grounding), 3) Human oversight (who reviews, authority, procedural rules), 4) Governance & redress (organisational policy, incident handling, legal/regulatory interfaces).
    • Three cross-cutting properties: traceability (audit trails, evidence), role clarity (named accountable actors across provider/integrator/deployer/overseer), continuous monitoring (runtime detection, change-control).
  • Mechanism families and maturity

    • Technical controls: detection (sample dispersion, alignment with references), grounding (retrieval-augmented generation), guardrails, automated verification — generally the most mature.
    • Human oversight: frequently invoked but under-specified (which humans, what evidence, authority to overrule) — low maturity.
    • Organisational governance: risk management, change control, incident response — emerging maturity with regulatory anchors.
    • Documentation & traceability: provenance records, logging, model cards/audit trails — moderate maturity but inconsistent standards.
    • Each family is assessed against maturity and regulatory anchors (EU AI Act, NIST, ISO).
  • Regulatory mapping

    • Mapped literature/mechanisms onto EU AI Act (high-risk obligations in force since 2 Aug 2026), NIST AI RMF GenAI Profile, ISO/IEC 42001, and sectoral guidance (healthcare, consumer finance, education, public sector).
    • Identified areas where existing literature supports regulatory requirements and where gaps remain.
  • Persistent gaps and tensions

    • Four persistent gaps: under-specified human oversight; absence of shared accountability metrics; disciplinary disconnection between technical and governance literatures; limited empirical evaluation of mechanisms.
    • Five structural tensions (not fully resolved by surveyed instruments), e.g., between model updates (dynamic systems) and the need for stable auditability; between provider vs deployer responsibilities; between transparency and IP/privacy; between standardized metrics vs domain-specific needs; between automated detection and human adjudication.
  • Hallucinations and formal constraints

    • Argues hallucination is systemic to probabilistic LLM generation: formal lower bounds (e.g., single-occurrence facts) and computability results imply hallucination cannot be eliminated—thus must be governed.
    • Detection formalisms discussed: dispersion/entropy across multiple samples, alignment score to external references with calibrated thresholds (τ), and combinations; threshold calibration is an accountability decision (trade-off between false positives/negatives).
  • Output

    • Proposes LAAF as an integrated accountability architecture, aligned to cybersecurity concerns (OWASP LLM Top 10), synthesising evidence rather than presenting a validated artefact.

Data & Methods

  • Review design

    • PRISMA-inspired systematic review covering Jan 2022–Mar 2026.
    • Search across five academic/technical databases (authors’ exact list not in this summary but the paper documents them).
    • Inclusion: studies addressing accountability mechanisms, conceptual treatments, technical detection/mitigation, governance, and regulatory instruments relevant to LLM deployments.
    • Result: 4,512 identified records → 122 primary studies included + 12 regulatory/standards documents analysed as primary sources.
  • Analytical approach

    • Conceptual synthesis: consolidated accountability into five dimensions and mapped failure modes (Section 5 of paper).
    • Mechanism synthesis: grouped mechanisms into four families and assessed maturity, regulatory anchors, and sectoral applicability (Section 6).
    • Regulatory mapping: comparative mapping to EU AI Act, NIST AI RMF (GenAI Profile), ISO/IEC 42001, and sectoral guidance (Section 7).
    • Identification of persistent gaps and tensions via cross-corpus analysis (Section 8).
    • Consolidation into LAAF: integrated the above into layered architecture (Section 9); explicitly stated as a synthesis (not validated empirically).
  • Formal/theoretical support

    • Incorporated formal results about inevitability of some hallucinations (references in paper) and probabilistic detection formalisms (entropy/dispersion, alignment scores, threshold calibration).

Implications for AI Economics

  • Liability, risk allocation, and incentives

    • Clear role-clarity requirements (provider vs integrator vs deployer) will shift where economic liability and compliance costs sit. Firms may re-arrange contracts, warranties, and terms of service to allocate legal and financial risk.
    • Under-specified human oversight creates regulatory and contractual uncertainty; economic actors will internalise this through higher premiums, stricter contracting, or avoidance of high-risk deployments.
  • Compliance costs and barriers to entry

    • Documentation, continuous monitoring, provenance logging, and certification to meet EU AI Act and similar standards impose fixed and ongoing costs. These create economies of scale favoring larger incumbents and can raise barriers to market entry for smaller firms/startups.
    • Conversely, specialised compliance-as-a-service providers and certification markets may emerge, creating new economic niches.
  • Insurance and risk-pooling markets

    • Absence of shared accountability metrics complicates insurance underwriting for LLM-caused harms. Development of standard metrics (e.g., calibrated hallucination thresholds, demonstrable oversight processes) would reduce information asymmetry and enable scalable insurance products.
    • Until such metrics exist, insurance costs will be high or cover limited perils; insurers may require demonstrable adherence to architectures like LAAF, certified monitoring, and documented human-in-loop procedures.
  • Product design and competitive strategy

    • Providers may differentiate by offering stronger provenance, explainability, and certified alignment tooling—creating product-market segmentation (safety/compliance premium).
    • Deployers (e.g., hospitals, banks) will trade off model capability vs. interpretability and auditability; high-regulation sectors will prefer integrations that maximise traceability even at some performance cost.
  • Dynamic update externalities and platform effects

    • The dynamic nature of LLMs (updates to models, retrieval corpora, prompts) produces externalities: a provider update may change downstream deployer risk. Contracts and platform governance will need mechanisms for notification, rollback, and coordinated change-control—economic frictions that can slow innovation or induce conservative update policies.
    • Large platform providers may internalise these coordination costs better than small vendors, reinforcing winner-take-most dynamics.
  • Measurement, metrics, and market signalling

    • Calibrated thresholds and measurable alignment scores become economic signals of trustworthiness. Firms able to credibly report standardised metrics can command higher prices or lower capital/insurance costs.
    • The current lack of shared metrics is a market failure: private actors have weak incentives to invest early in costly, non-excludable metrics. Regulatory intervention or industry consortia may be necessary to create public goods (standard metrics, benchmark datasets).
  • Sectoral heterogeneity and social welfare

    • Sector-specific guidance (healthcare, finance, education, public sector) implies differential compliance burdens and welfare implications. In health and finance, the social cost of errors is high—leading to stricter oversight and higher costs that may reduce adoption speed but increase safety. In less regulated sectors, faster deployment may yield consumer surplus but higher systemic risk.
    • Policymakers must weigh the trade-off between innovation diffusion and the social cost of harms when calibrating sectoral regulatory strictness.
  • Research & policy priorities for economics

    • Prioritise development of shared accountability metrics and standard benchmarking that are interoperable across sectors—this reduces transaction costs and insurance friction.
    • Study how compliance costs affect market structure, especially entry and concentration, and design policy measures (e.g., subsidies, standardisation support) to mitigate anti-competitive effects.
    • Explore contract designs and liability rules (e.g., mandatory disclosure, strict liability for deployers in certain contexts) that align incentives across providers, integrators, and deployers.
    • Model the equilibrium effects of mandatory provenance/traceability requirements on innovation rates, platform dominance, and consumer welfare.
  • Practical takeaways for economic actors

    • Firms deploying LLMs should anticipate and budget for ongoing monitoring, provenance logging, and human-in-loop procedures; these are not one-time compliance tasks.
    • Early investment in traceability and standardised metrics can yield lower insurance and regulatory friction costs later, and can be a credible market signal.
    • Smaller actors should consider partnering with compliance-specialist providers or using certified platforms to manage the fixed costs of compliance.

Note: LAAF is presented as a synthesis of surveyed evidence rather than a validated operational artefact. The paper emphasises the need for empirical evaluation and shared metrics—both essential for translating the architecture into workable economic incentives and market mechanisms.

Assessment

Paper Typereview_meta Evidence Strengthmedium — The review is comprehensive (5 databases, 4,512 records, 122 primary studies plus 12 regulatory/standards documents) and synthesises conceptual, technical, and regulatory literature, but the underlying corpus is heavily conceptual with limited empirical evaluation and few causal or quantitative studies validating proposed mechanisms, which reduces the overall strength of empirical claims. Methods Rigorhigh — The authors report a PRISMA-inspired systematic search, explicit review questions, a large screened set and a clearly described synthesis (taxonomies, maturity assessments, regulatory mapping). However, the text is 'PRISMA-inspired' rather than demonstrating full PRISMA registration or all reproducibility details in the excerpt (e.g., full eligibility criteria, quality appraisal metrics and extraction protocol are not shown), so some methodological transparency details are uncertain. SampleSystematic literature corpus covering Jan 2022–Mar 2026 identified from five databases (4,512 records screened), with 122 primary studies included alongside 12 regulatory and standards documents (EU AI Act, NIST AI RMF GenAI Profile, ISO/IEC 42001, sectoral guidance from FDA, EMA, CFPB, EBA, UNESCO, US Dept. of Education, OECD, etc.). The included studies span computer science (technical controls, hallucination detection), governance/regulatory analyses, ethics, and sectoral guidance; many are conceptual or technical evaluations rather than large-scale empirical causal studies. Themesgovernance org_design adoption human_ai_collab GeneralizabilityTime-bounded coverage (Jan 2022–Mar 2026); LLM capabilities and regulatory environments evolve rapidly, so findings may age quickly, Geographic/regulatory focus biased toward EU, US (NIST), and international standards (ISO); less coverage of other regulatory regimes, Sectoral mapping emphasizes healthcare, consumer finance, education, and public sector—other industries may face different accountability needs, Underlying literature is largely conceptual/technical with limited empirical validation, so practical applicability of some mechanisms is uncertain, Possible language/selection biases and 'PRISMA-inspired' methodology may limit reproducibility of exact corpus

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The review identified 4,512 records and included 122 primary studies, along with 12 regulatory and standards documents analyzed as primary sources. Governance And Regulation positive Size and composition of the reviewed evidence corpus
Reading fidelity high
Study strength medium
n=4512
0.24
The reviewed literature has four persistent accountability gaps: under-specified human oversight, a lack of shared accountability metrics, disconnection between technical and governance disciplines, and limited empirical evaluation of accountability mechanisms. Governance And Regulation negative Completeness and empirical maturity of accountability practices
Reading fidelity high
Study strength medium
n=122
0.24
The review concludes that accountability for LLM applications should be treated as both a lifecycle property and a distributed property allocated across named actors. Governance And Regulation positive Clarity and continuity of responsibility assignment across LLM deployment
Reading fidelity high
Study strength medium
n=122
0.24
The surveyed literature does not support a defensible pooled estimate of hallucination frequency because the studies use incompatible definitions, prompts, and domains. Error Rate null_result Pooled hallucination frequency
Reading fidelity high
Study strength medium
not reported
0.24
The review characterizes hallucination as a systemic property of probabilistic generation rather than merely a transient defect of particular model releases. Error Rate negative Reliability and factuality of LLM outputs
Reading fidelity high
Study strength medium
not reported
0.24
Hallucination-detection approaches based on response dispersion can provide sample-level evidence of hallucination without access to model internals, but they cannot detect confident, consistent errors. Error Rate mixed Hallucination detection capability
Reading fidelity high
Study strength medium
n=10
0.24
The proposed Layered Accountability Architecture Framework (LAAF) is a synthesis of the reviewed evidence and has not been validated as an artefact. Governance And Regulation positive Integrated accountability architecture for LLM applications
Reading fidelity high
Study strength speculative
n=122
0.04

Notes