The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Internal audits, if redesigned for frontier AI, can give boards and regulators stronger, system-wide assurance over catastrophic risks; practical trade-offs—scope (model, system, governance), sourcing, cadence and access to sensitive information—determine effectiveness and must be balanced against security and organizational constraints.

How frontier AI companies could implement an internal audit function
Francesca Gomez, Adam Buick, Leah Ferentinos, Haelee Kim, Elley Lee · December 16, 2025
arxiv commentary n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Francesca Gomez unresolved corpus identity
  2. Adam Buick unresolved corpus identity
  3. Leah Ferentinos unresolved corpus identity
  4. Haelee Kim unresolved corpus identity
  5. Elley Lee unresolved corpus identity

Semantic Scholar

Latest observation:

  1. F. Gómez provider ID
  2. A. Buick provider ID
  3. Leah Ferentinos provider ID
  4. Haelee Kim provider ID
  5. Elley Lee provider ID
The paper proposes a framework for internal audit functions tailored to frontier AI developers, analyzing audit scope, sourcing, cadence, and access trade-offs and arguing that well-designed internal audit can deliver system-wide assurance over catastrophic risk controls.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention has been paid to how internal organizational assurance should be structured to provide sustained, evidence-based oversight of catastrophic and systemic risks. This paper examines how an internal audit function could be designed to provide meaningful assurance for frontier AI developers, and the practical trade-offs that shape its effectiveness. Drawing on professional internal auditing standards, risk-based assurance theory, and emerging frontier-AI governance literature, we analyze four core design dimensions: (i) audit scope across model-level, system-level, and governance-level controls; (ii) sourcing arrangements (in-house, co-sourced, and outsourced); (iii) audit frequency and cadence; and (iv) access to sensitive information required for credible assurance. For each dimension, we define the relevant option space, assess benefits and limitations, and identify key organizational and security trade-offs. Our findings suggest that internal audit, if deliberately designed for the frontier AI context, can play a central role in strengthening safety governance, complementing external evaluations, and providing boards and regulators with higher-confidence, system-wide assurance over catastrophic risk controls.

Summary

Main Finding

An internal audit function—deliberately designed for frontier-AI developers—can provide durable, board-level, evidence-backed assurance over catastrophic and systemic AI risks. The most effective design integrates multi-level audits (model, system, governance), uses a hybrid sourcing model anchored by an internal Chief Audit Executive (CAE) with external specialists, differentiates audit frequency by risk, and implements tiered/controlled information-access regimes to balance assurance quality with information-security concerns.

Key Points

  • Purpose and relevance

    • Internal audit = independent, board‑reporting function that provides enterprise-wide assurance over governance, risk management, and controls.
    • Relevant to regulatory obligations (e.g., EU General-Purpose AI Code of Practice commitment 8) that require board-level assurance over systemic risk management.
    • Complements existing external technical evaluations and security certifications by covering organizational controls and pre-deployment internal use.
  • Four core design dimensions

  • Audit scope: model-level, system-level, governance-level (or integrated multi-level) - Model-level: most direct view of dangerous capabilities but high cost, high deployment friction, rapid assurance decay. - System-level: checks containment/enforcement (logging, access controls, monitoring); more stable and lower marginal technical expertise needed but cannot quantify true model capability. - Governance-level: checks decision-rights, escalation, board oversight; low cost, low friction, temporally stable but can be disconnected from technical reality. - Multi-level integration yields the most resilient assurance (detects/compensates for degradation in any single layer) but is costliest.
  • Sourcing: in-house; co-sourced (internal CAE + external specialists); fully outsourced - Hybrid/co-sourcing anchored by an internal CAE is recommended: preserves board reporting and independence while allowing specialist skills and external credibility. - Fully outsourced maximizes perceived independence but raises security and cultural risks; fully internal reduces security exposure but faces independence/perception challenges.
  • Frequency and cadence: annual; semi-annual/quarterly; ad-hoc/rapid; continuous - Model-level assurance requires higher cadence due to rapid capability change; system/governance audits can be less frequent. - Rolling/continuous planning and ad-hoc rapid reviews are important given fast iteration and episodic risk spikes.
  • Information access: structural, procedural, operational, technical - High-quality assurance, particularly for model- and system-level audits, requires access to sensitive technical artifacts (evaluations, logs, code, models). - Access raises information-security risks; mitigation options include tiered access, clean-room environments, time-bounded and rate-limited access.

  • Implementation trade-offs

    • Assurance quality vs security exposure: more access improves assurance but creates attack surface and leakage risk.
    • Independence vs insider knowledge: purely external auditors may lack necessary context; internal auditors may face perceived or real conflicts.
    • Assurance relevance vs operational friction: higher-frequency, technically deep audits can slow development and add cost.
    • Firms can phase capability: start with governance/system audits and build toward integrated model-level capability as processes and secure information-handling mature.
  • Practical mechanisms discussed

    • Risk-based audit planning (audit universe, CAE proposes board‑approved plan).
    • Tiered access models, clean rooms, time-limited credentials to manage sensitive artifact exposure.
    • Differentiated evidence expectations across lifecycle stages (appendix D).

Data & Methods

  • Nature of the paper: conceptual and normative analysis rather than empirical evaluation.
  • Sources and frameworks drawn upon:
    • Professional internal auditing standards and guidance (Institute of Internal Auditors, IIA).
    • Risk-based assurance theory and established audit practices (IT/cyber/audit frameworks such as COBIT, NIST Cybersecurity Framework).
    • Emerging frontier‑AI governance literature, public developer disclosures, and recent capability assessments (e.g., METR, developer whitepapers).
  • Analytical approach:
    • Map conventional internal-audit methodologies and audit types to frontier-AI risk domains.
    • Define option spaces for each design dimension and systematically assess benefits, limitations, and organizational/security trade-offs.
    • Produce prescriptive design guidance and illustrative mechanisms (appendices detail stages of risk-based audit planning, planning cadence options, and access-management techniques).
  • Empirical content: no new empirical datasets or experiments; relies on literature synthesis, standards mapping, and scenario-based reasoning.

Implications for AI Economics

  • Internal audit as risk-management capital

    • Adds an organizational governance cost (set-up and running costs for CAE, specialist evaluators, secure infrastructure).
    • Produces value by reducing residual systemic risk, improving regulatory compliance, and increasing board/regulator/investor confidence.
    • Should be treated as a form of risk-management investment that can alter expected downside tail risks and thus firm valuations and cost of capital.
  • Market structure and competitive dynamics

    • Fixed costs and scarce specialist skills (secure evaluation infrastructure, vetted auditors) can raise barriers to entry; incumbents or well‑capitalized firms may adopt more comprehensive audit systems faster.
    • Conversely, credible internal audit and assurance could be a differentiator, enabling safer public-facing releases and partnerships, and potentially reaping commercial and reputational benefits.
  • Pace-of-deployment and race effects

    • High-friction, frequent model-level audits can slow deployment, which may reduce “race” incentives and expected probability of catastrophic outcomes—an important externality for social-welfare analysis.
    • Policy design should consider incentives: if regulators or markets reward audited assurance, firms may internalize safety benefits and choose safer pacing.
  • Insurance, liability, and capital allocation

    • Strong internal audit practices improve observability and evidence for insurers, potentially enabling insurance products for certain AI risks and lowering premiums for compliant firms.
    • Audits could also affect liability assessments and contract terms (e.g., procurement, third-party use), altering downstream market behavior and costs.
  • Regulatory and policy implications

    • Regulatory standards that recognize internal audit as a compliance mechanism (as in EU Code of Practice) can shift enforcement toward organizational assurance rather than only ex-post sanctions.
    • Economists modeling AI policy should incorporate the heterogeneity of firm assurance capacity and the costs/benefits of different audit designs when evaluating regulation (e.g., scalability of audit requirements, subsidies for SMEs, standardized evidence frameworks).
  • Signaling and information externalities

    • Boards’ and firms’ ability to credibly signal effective systemic-risk controls depends on audit independence and credible handling of sensitive evidence; hybrid models anchored by an internal CAE with vetted external specialists can increase signal credibility without undermining security.
    • Standardized assurance frameworks would reduce information asymmetries between firms, investors, and regulators.

Practical considerations for economists and policymakers - In cost–benefit assessments of safety regulation, include internal-audit set-up and operating costs, and model how audit-induced deployment delays change risk trajectories. - Consider incentives/subsidies for smaller developers to avoid concentration of frontier capabilities driven by compliance costs. - Promote standards for evidence handling and auditor accreditation to reduce information-security risk while preserving credibility of assurance.

Assessment

Paper Typecommentary Evidence Strengthn/a — Conceptual and normative analysis drawing on professional standards and governance literature rather than empirical or causal evidence; no primary data or causal identification attempted. Methods Rigormedium — Systematic synthesis of internal auditing standards, risk-based assurance theory, and emerging frontier-AI governance literature provides a disciplined conceptual framework, but the paper lacks empirical validation, pre-registered protocols, or case-study triangulation to test practical effectiveness. SampleNo empirical sample; the paper is a conceptual analysis synthesizing professional internal auditing standards, risk-based assurance theory, and frontier-AI governance literature to develop design dimensions for internal audit functions. Themesgovernance org_design adoption GeneralizabilityFocused on frontier AI developers; recommendations may not transfer to smaller or non-frontier firms, Applicability varies across regulatory jurisdictions and legal regimes, Relies on firms' willingness to grant auditors access to highly sensitive models and data, which may be constrained by security concerns, Effectiveness depends on availability of personnel with both deep technical ML knowledge and audit expertise, Rapid technical and threat evolution could outdate specific control prescriptions

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. Governance And Regulation mixed operating environment characterized by technical progress, risk exposure, and regulatory scrutiny
Reading fidelity high
Study strength speculative
not reported
0.01
Comparatively little attention has been paid to how internal organizational assurance should be structured to provide sustained, evidence-based oversight of catastrophic and systemic risks. Governance And Regulation negative degree of attention to internal organizational assurance for catastrophic/systemic AI risks
Reading fidelity high
Study strength speculative
not reported
0.01
An internal audit function could be designed to provide meaningful assurance for frontier AI developers, subject to practical trade-offs that shape its effectiveness. Governance And Regulation positive meaningfulness/effectiveness of assurance provided by an internal audit function
Reading fidelity high
Study strength speculative
not reported
0.01
Four core design dimensions for internal audit in the frontier AI context are: (i) audit scope across model-level, system-level, and governance-level controls; (ii) sourcing arrangements (in-house, co-sourced, and outsourced); (iii) audit frequency and cadence; and (iv) access to sensitive information required for credible assurance. Governance And Regulation mixed relevant design dimensions of internal audit functions
Reading fidelity high
Study strength low
not reported
0.03
For each design dimension, there are identifiable benefits and limitations, and key organizational and security trade-offs that affect audit effectiveness. Governance And Regulation mixed benefits, limitations, and trade-offs associated with internal audit design choices
Reading fidelity high
Study strength low
not reported
0.03
Internal audit, if deliberately designed for the frontier AI context, can play a central role in strengthening safety governance. Governance And Regulation positive strength of safety governance
Reading fidelity high
Study strength speculative
not reported
0.01
Internal audit can complement external evaluations. Governance And Regulation positive relationship between internal audit and external evaluation functions
Reading fidelity high
Study strength speculative
not reported
0.01
Internal audit can provide boards and regulators with higher-confidence, system-wide assurance over catastrophic risk controls. Governance And Regulation positive confidence of boards and regulators in system-wide assurance over catastrophic risk controls
Reading fidelity high
Study strength speculative
not reported
0.01

Notes