The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Frontier AI labs show the same organizational drift that preceded major technological disasters: safety procedures can be completed in form while organizational dynamics normalize dangerous shortcuts. Independent oversight, structural separation of safety functions, and cultural reforms are needed to prevent incremental drift from becoming catastrophe.

The Normalization of Deviance in AI Development
Emilio Barkett, Alexander Kimpton, Daniel Graham, Yusuf Kundgol · September 04, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Emilio Barkett unresolved corpus identity
  2. Alexander Kimpton unresolved corpus identity
  3. Daniel Graham unresolved corpus identity
  4. Yusuf Kundgol unresolved corpus identity
AI development organizations are susceptible to the normalization of deviance—production pressure, false reassurance from prior successes, structural secrecy, and eroding independent oversight can convert rule-following safety processes into a pathway toward catastrophic failures.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the institutions building these systems are themselves predisposed to drift toward failure. This paper argues that they are. Regardless of how capable AI systems become, the organizations building them face the same structural dynamics that preceded past major technological disasters. Drawing on case studies of the Space Shuttle Challenger, the Three Mile Island accident, and the Boeing 737 MAX crashes, this paper identifies the common structural mechanisms preceding each failure and maps them onto contemporary AI development. The findings suggest that existing safety infrastructure may provide less protection than it appears, as organizations can complete safety processes in full compliance and still produce catastrophic outcomes. The pre-disaster period of AI development is still underway; the purpose of this paper is to make these dynamics legible while they can still be interrupted.

Summary

Main Finding

Organizations building frontier AI systems are structurally prone to “normalization of deviance” — an incremental drift in which practices that violate safety norms become treated as acceptable because they repeatedly produce no immediate catastrophe. This organizational dynamic (driven by production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight) can make existing safety processes ineffective: firms can formally complete safety checks yet still drift toward catastrophic outcomes. The pre-disaster period for AI is ongoing, and interrupting these dynamics requires structural, cultural, governance, and technical changes.

Key Points

  • The paper reframes AI risk from an individual-system capability problem to an organizational problem: failure can arise from ordinary organizational dynamics, not only from more powerful models.
  • Four recurring mechanisms that underlie normalization of deviance:
    • Production pressure: competitive, reputational, and geopolitical incentives make safety work look like friction and shift the burden of proof onto caution.
    • False assurance from prior success: passing benchmarks and benign deployments get treated as proof of safety, raising the implicit baseline.
    • Structural secrecy: division of labor and hierarchical incentives prevent safety-relevant information from influencing deployment decisions.
    • Erosion of independent oversight: expertise concentrates within developers and self-certification becomes the default, reducing meaningful external checks.
  • Historical analogues: Space Shuttle Challenger (O-ring erosion and burden-of-proof reversal), Three Mile Island (informal workarounds and opacity), Boeing 737 MAX (self-certification and regulatory capture). These cases illustrate how the four mechanisms operate and compound.
  • Evidence of these dynamics in AI:
    • Intense commercial and strategic pressure among frontier labs (including US–China rivalry).
    • Benchmark-driven assessments treated as safety evidence despite known blind spots.
    • Safety teams, red-teams, and product teams separated by incentives and reporting lines; safety staff departures publicly reported.
    • Governance dominated by developer-led standards, self-evaluation, and limited independent access — a de facto self-certification regime.
  • Recommendations (high level): create structural independence for safety functions; require pre-mortems and formal escalation for safety objections; build epistemic cultures that treat past deployments as precedent, not proof; establish third-party independent evaluation and incident-reporting systems; fund interpretability and mandate staged deployments with enforceable rollback criteria.

Data & Methods

  • The study is qualitative and comparative, drawing on:
    • A theoretical framework synthesized from organizational failure literature (Merton, Turner, Hughes, Perrow, Sagan, Vaughan, Dekker) centered on normalization of deviance.
    • Three well-documented historical case studies (Challenger, Three Mile Island, Boeing 737 MAX) selected for rich investigatory records that reveal organizational dynamics prior to catastrophe.
    • Mapping and analogical reasoning that projects the structural mechanisms identified in those historical cases onto contemporary frontier AI development, supported by public reporting (e.g., incidents, safety-team departures, governance practices, benchmark histories) and recent examples (e.g., July 2026 Hugging Face incident referenced by authors).
  • No original quantitative dataset or econometric analysis; instead, the method is literature synthesis, case-comparison, and institutional mapping.
  • The paper acknowledges limits to historical analogy and calls for empirical extensions (e.g., direct organizational studies, incident data collection, measurement of reporting flows and oversight efficacy).

Implications for AI Economics

  • Incentives and market failures
    • Safety as a negative externality: firms capture benefits of rapid deployment while diffusing broader societal risks, creating classic under-provision of safety.
    • Short-term competition and geopolitical framing intensify “race” incentives, skewing firm-level cost–benefit calculations against costly, slow safety work even when social welfare would favor it.
    • Benchmark- and deployment-based legitimacy can create perverse incentives: firms optimizing for benchmark scores may underinvest in unmeasured safety dimensions.
  • Regulation, certification, and liability
    • Self-certification is fragile: when regulatory assessment depends on developer-provided expertise, markets can underprice systemic risk and insurance markets may misestimate tail exposures.
    • Independent third-party evaluation and transparent incident reporting can reduce information asymmetries, improve social pricing of risk (e.g., through insurance premiums, liability exposure), and align private incentives with public safety.
    • Mandated structural independence for safety functions and enforceable staged-deployment rules change firm-level constraint sets and raise the private cost of risky deployment, altering equilibrium behavior.
  • Investment and R&D allocation
    • Public and private funding priorities should shift to include interpretability, red-team transparency, and governance research; subsidizing these lowers the private cost of safer development and corrects R&D market failures.
    • Investors and corporate governance should internalize long-term systemic risk: due diligence ought to include organizational safety structures, independence of safety teams, and histories of transparency and incident reporting.
  • Insurance, risk pooling, and systemic externalities
    • Without transparent, independent assessments, insurers cannot properly price systemic AI risks, possibly leading to market withdrawal or prohibitively high premiums for high-capability deployments.
    • Public risk-pooling or backstop mechanisms (analogous to deposit insurance, nuclear liability frameworks, or aviation oversight funds) may be required to internalize systemic tail risks and prevent underinsurance.
  • Labor and knowledge concentration
    • Expertise concentration in frontier labs creates an epistemic monopoly that hinders credible external oversight; policies that broaden technical capacity across academia, regulators, and independent labs reduce asymmetries and improve market functioning.
  • Policy trade-offs
    • Stronger oversight and slower deployment carry opportunity costs (lost first-mover rents, delayed product-market benefits). Policymakers must balance innovation incentives with precaution, potentially via staged and conditional licensing, adaptive regulation, and compensatory incentives (grants, safe-harbor provisions for transparent reporting).
  • Measurement and monitoring for economic models
    • For economists modeling AI externalities, the paper highlights key organizational variables to incorporate: production-pressure intensity, transparency indices, independence of safety reporting, and degree of self-certification. These variables can mediate the translation of technical capabilities into realized systemic risk and should inform welfare analyses, regulatory cost–benefit calculations, and insurance models.

Overall, the paper implies that correcting organizational incentives and governance failures is as important as bounding model capabilities: economic policy tools (regulation, accountability, subsidies, insurance design) should target the organizational mechanisms that allow deviance to normalize, not only technical safety features.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is a conceptual and historical-analogical analysis rather than an empirical causal study; it synthesizes case studies and published reports to build a theoretical argument, so standard strength-of-evidence metrics for causal inference do not apply. Methods Rigormedium — The authors draw on well-documented, exhaustively investigated historical cases and relevant AI incidents and literature, and they clearly map identified mechanisms onto AI development; however, the approach is analogical and descriptive without systematic empirical testing, formal counterfactuals, or new primary data to validate the frequency or magnitude of the mechanisms in contemporary AI labs. SampleNo original quantitative sample; the paper synthesizes organizational-failure literature and detailed historical case studies (Space Shuttle Challenger, Three Mile Island, Boeing 737 MAX), official investigations and reports (e.g., Presidential and Congressional inquiries), documented incidents in AI (cited example: July 2026 Hugging Face incident), public departures and statements from frontier labs, and contemporary AI governance literature. Themesgovernance org_design GeneralizabilityRelies on historical analogies—AI development (software, ML models) has important differences from the physical-engineering contexts examined., Case selection focuses on high-profile disasters; may over-emphasize mechanisms that are salient in spectacular failures versus routine practice., No systematic, cross-organizational empirical validation of how widespread the mechanisms are across AI labs or geographies., Institutional and regulatory environments vary by country and firm; recommendations may not transfer uniformly., Technological differences (e.g., opacity of models, rapid iteration cycles) could both amplify and alter mechanisms in ways not fully captured by the analogies.

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A systematic review of 33 studies found consistent evidence of production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight across oil and gas, nuclear, aviation, healthcare, and rail industries. Organizational Efficiency negative Presence of organizational mechanisms associated with safety failures
Reading fidelity high
Study strength medium
n=33
0.12
Normalization of deviance is an incremental organizational process in which practices that violate safety norms become redefined as normal after repeated non-disastrous outcomes. Ai Safety And Ethics negative Organizational normalization of unsafe practices
Reading fidelity high
Study strength medium
not reported
0.12
The Challenger launch decision was preceded by repeated observations of O-ring erosion that were retrospectively interpreted as evidence of safety because earlier flights had not produced a catastrophe. Decision Quality negative Safety decision quality and launch failure
Reading fidelity high
Study strength high
n=7
0.2
The Three Mile Island accident was enabled in part by informal workarounds and accumulated complexity that created a gap between formal procedures and actual operating practice. Decision Quality negative Operator decision quality and nuclear safety
Reading fidelity high
Study strength high
not reported
0.2
The Lion Air 610 and Ethiopian Airlines 302 crashes were caused by the MCAS software system repeatedly commanding the aircraft nose down in response to faulty sensor data. Error Rate negative Aircraft safety and accident occurrence
Reading fidelity high
Study strength high
n=2
0.2
Boeing’s decision to modify the existing 737 rather than design a new aircraft created a trajectory in which engineering compromises became increasingly difficult to reverse. Decision Quality negative Engineering and safety decision quality
Reading fidelity high
Study strength high
not reported
0.2
The FAA’s use of Boeing employees as Authorized Representatives contributed to a form of self-certification in which the organization being regulated performed much of the safety analysis and certification work. Governance And Regulation negative Independence and effectiveness of regulatory oversight
Reading fidelity high
Study strength high
not reported
0.2
The paper argues that AI development currently combines production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight. Ai Safety And Ethics negative Organizational exposure to conditions associated with AI safety failure
Reading fidelity high
Study strength low
not reported
0.06
In AI development, prior successful deployments and benchmark performance can be treated as safety evidence, progressively making it harder to justify stricter standards as capabilities increase. Ai Safety And Ethics negative Validity of AI safety assurance and evaluation standards
Reading fidelity high
Study strength medium
not reported
0.12
The paper argues that safety-relevant information in frontier AI organizations may fail to reach people with authority to act because safety, product, and leadership teams operate in different organizational and incentive environments. Organizational Efficiency negative Flow of safety information and organizational safety decision quality
Reading fidelity high
Study strength low
not reported
0.06
The paper argues that current frontier AI governance commonly relies on self-certification because responsible scaling policies and model evaluations are defined, conducted, and enforced by the same organizations making deployment decisions. Governance And Regulation negative Independence and effectiveness of AI safety oversight
Reading fidelity high
Study strength low
not reported
0.06
The paper recommends structurally independent AI safety organizations with authority to delay or halt deployment and with reporting channels outside the production management chain. Governance And Regulation positive AI safety oversight and prevention of unsafe deployment
Reading fidelity high
Study strength speculative
not reported
0.02
The paper recommends independent third-party AI evaluation infrastructure that does not depend on developer cooperation, is funded independently, and can publish its own assessments of safety-relevant model properties. Governance And Regulation positive Independence and effectiveness of AI safety evaluation
Reading fidelity high
Study strength speculative
not reported
0.02

Notes