The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new CASE framework argues enterprises must treat agentic AI as four distinct governance problems — and finds they do not: public evidence shows multi-layer failure trajectories are common, no analyzed tooling fully covers emergence risks, and every scored public deployment fell into the lowest maturity band.

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI
Srinivas Telukunta, Georgios Nektarios Lilis, Lucio Baron · August 10, 2026
arxiv theoretical medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Srinivas Telukunta unresolved corpus identity
  2. Georgios Nektarios Lilis unresolved corpus identity
  3. Lucio Baron unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Srinivas Telukunta provider ID
  2. Georgios Nektarios Lilis provider ID
  3. Lucio Baron provider ID
The CASE framework maps four established sciences to four scales of agentic AI governance, shows that emergent multi-agent behavior (Layer 2) is systematically under-governed (the 'Emergence Gap'), and validates this with public evidence that most failures cross layers while tooling and deployments lack Layer 2 coverage and maturity.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single-agent assurance non-compositional), Supervisory cybernetics to human-agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross-layer coupling conditions, including a zero-touch deployment paradox where excellence at one-layer strains the others, and trace twenty-plus enterprise controls to their classical constructs. Three empirical studies validate the thesis: 82 percent of documented production agent failures are multi-layer trajectories; none of 22 ecosystem tools offers full Layer 2 (emergence) coverage; and all 35 scored public deployments fall in the lowest maturity band. We name this mismatch, risk realized at the emergence layer against capability barely offered and practice absent, the Emergence Gap. A five-level maturity model with a non-compensatory bottleneck-weighted index and assessment instrument operationalizes CASE as a scientific rather than process maturity model, grounded in production enterprise agentic platforms. As EU AI Act Article 14 makes effective human oversight a legal requirement, only architectures satisfying requisite variety can make oversight real rather than ceremonial.

Summary

Main Finding

The paper introduces CASE, a four-layer governance architecture (Control theory, complex Adaptive systems, Supervisory cybernetics, Engineering operations) that maps a distinct governing science to each scale of agentic AI (individual agent, agent collectives, human–agent teams, automated fleets). CASE shows that single-discipline governance is insufficient: failures are usually multi-layer trajectories, cross-layer coupling creates new governance paradoxes (notably a "zero-touch deployment paradox") and there is a widespread Emergence Gap — empirical and tooling evidence that Layer 2 (emergent multi-agent behavior) is under-governed. The authors operationalize CASE with formal conditions (including a testable requisite‑variety oversight inequality), a non‑compensatory five‑level maturity model (bottleneck-weighted index), and an implementable mechanism inventory tied to classical constructs.

Key Points

  • CASE mapping:
    • L1 (individual agent) → Control theory: observability, controllability, stability; guardrails as feedback; observer telemetry as the state estimator.
    • L2 (agent collectives) → Complex adaptive systems: emergence, criticality, cascades, monoculture risk, shared-state feedback.
    • L3 (human–agent teams) → Supervisory cybernetics: Law of Requisite Variety, Viable System Model, algedonic channels for bypass escalation.
    • L4 (enterprise fleet) → Engineering operations (AgentOps/SRE): decision‑quality SLOs, error budgets extended to autonomy, escalation and rollback procedures.
  • Novel formal contributions:
    • Converts requisite variety into an engineering inequality that can be tested and audited (human oversight capacity × engineered amplification ≥ peak agent behavioral variety).
    • Treats autonomy as a controlled variable by closing the loop between decision‑quality error budgets and autonomy modulation.
    • Derives cross‑layer coupling conditions, including the zero‑touch paradox: greater automation and deployment velocity (Layer 4 excellence) mechanically increases interaction variety and strains oversight (Layers 2 and 3).
    • Presents a non‑compensatory maturity index in which the weakest layer dominates; funding the binding (weakest) layer is the optimal marginal investment.
  • Mechanism inventory links 20+ deployable enterprise controls to classical constructs (e.g., receding‑horizon planning ↔ model predictive control; shared‑memory governance ↔ stigmergy; algedonic channel ↔ bypass escalation).
  • Grounding and validation:
    • Authors draw on experience operating production agentic platforms at Fortune‑50 scale.
    • Three empirical studies (web‑scale): failure taxonomy, tooling capability gap, maturity distribution.
    • Key empirical findings reported:
      • 82% of documented production agent failures are multi‑layer trajectories.
      • None of 22 evaluated ecosystem tools offers full Layer 2 (emergence) coverage.
      • All 35 publicly scored deployments fall in the lowest maturity band — the Emergence Gap manifest in practice.
  • Practical outputs: a five‑level CASE maturity model, a full assessment instrument (Appendix A), and a Zero Touch Agent Deployment reference architecture (Section 7).

Data & Methods

  • Validation design: three complementary web‑scale empirical studies, with coding protocols published in Appendix B:
  • Failure taxonomy mapping — coded documented production agent failures across public incident reports to identify layer attribution and cross‑layer trajectories (result: 82% multi‑layer).
  • Tooling capability analysis — surveyed 22 ecosystem tools/platforms for coverage across CASE layers, with special attention to Layer 2 (emergence) controls (result: no tool provided full Layer 2 coverage).
  • Enterprise maturity distribution — applied the CASE assessment instrument to 35 public deployments/platforms and scored maturity across layers using the non‑compensatory, bottleneck‑weighted index (result: all 35 in lowest maturity band).
  • Formalism and indices: the paper formalizes:
    • L1 as a discrete‑time control system with observer, controller, and safety (Lyapunov‑style) condition.
    • The oversight requisite‑variety inequality (convertible into measurable artifacts like variety budgets and amplification chains).
    • An autonomy modulation rule tying decision‑quality error budgets to allowed autonomy.
    • A composite maturity index (bottleneck weighting, geometric mean as robustness limit) and associated scoring bands (Equation references in paper).
  • Reproducibility: coding protocols and the full assessment instrument are made available in appendices to enable replication.

Implications for AI Economics

  • Investment prioritization and capital allocation:
    • The non‑compensatory maturity index implies diminishing returns from spreading governance investment evenly; firms maximize safety/economic value by prioritizing the binding (weakest) CASE layer first. This changes ROI calculations for governance spending and suggests targeted CAPEX/OPEX for Layer 2 and Layer 3 may yield higher marginal value than incremental Layer 4 tooling.
  • Market failures and business opportunities:
    • The Emergence Gap (lack of Layer 2 tooling and governance) is both a systemic risk and a commercial opportunity. Tool vendors and platform providers that can credibly deliver emergence‑aware controls (monitoring for cascades, percolation detection, shared‑state governance) will meet a large unmet market need; absence of such products risks underprovision of public‑good safety features and potential regulatory intervention.
  • Regulatory compliance costs and competitive dynamics:
    • EU AI Act Article 14 (effective human oversight requirement) becomes operationally meaningful via CASE: firms must demonstrate requisite‑variety chains (human + engineered amplification). Compliance will favor organizations with CASE‑aligned architectures and assessment artifacts (variety budgets, algedonic channels, maturity scores), raising barriers to entry for smaller firms or startups lacking these investments.
  • Labor and organization:
    • The Law of Requisite Variety implies that unaided human oversight scales poorly. Firms will need investments in augmentation (automation that amplifies human oversight variety) and in organizational redesign (recursive Viable System Model structures). This shifts labor demand toward roles that design oversight amplifiers, interpret algedonic signals, and manage cross‑layer coupling — higher‑skilled, potentially fewer staff but with different skills.
  • Operational cost structure and productivity:
    • Extending error budgets to decision quality creates new operational metrics (decision‑quality SLOs, autonomy budgets) that replace or supplement throughput/availability metrics. These metrics will affect pricing, SLA design, and cost accounting for agentic services (e.g., per‑decision quality tiers).
    • The zero‑touch deployment paradox warns that scaling automation without commensurate investment in Layers 2 & 3 can generate negative externalities (more incidents, higher mitigation cost), changing the calculus of full automation vs. staged autonomy.
  • Systemic risk and externalities:
    • Monoculture risk and shared‑state feedback in agent collectives create correlated failure modes across firms that use the same base models and shared tools. This increases systemic tail risk, with implications for insurers, regulators, and financial stability assessments. Economists and policymakers should treat Layer‑2 governance as a public‑good problem where private incentives underprovide resilience.
  • Measurement, auditing, and valuation:
    • CASE supplies measurable artifacts (assessment scores, variety budgets, algedonic channel traces) that can be audited, insured, and potentially required by regulators. This reduces information asymmetry in markets for agentic services and enables more precise risk pricing, M&A due diligence, and corporate governance assessments.
  • Policy prescriptions:
    • Regulators should require demonstrable requisite‑variety evidence and non‑compensatory maturity reporting rather than only process checklists.
    • Subsidies, standards, or certification programs targeting Layer 2 tooling and cross‑layer testing (chaos‑style experiments for emergence) could correct underinvestment and lower systemic risk.

In short, CASE reframes agentic AI governance as a multidimensional economic problem: the distribution of governance investments across layers, the presence of systemic externalities from insufficient Layer‑2 controls, and the measurable nature of requisite variety all change how firms, markets, and regulators should evaluate cost, risk, and value in the age of agentic AI.

Assessment

Paper Typetheoretical Evidence Strengthmedium — The paper is primarily a theoretical and conceptual contribution that is supported by three descriptive, web-scale empirical studies using public sources and documented coding protocols; these provide convergent, replicable evidence for the framework’s practical gaps (e.g., the Emergence Gap) but are non-experimental, subject to selection and publication bias, and do not establish causal effects on economic outcomes. Methods Rigormedium — The theoretical mapping and formalism are well-grounded in established literatures (control theory, complex systems, cybernetics, SRE) and the paper derives formal coupling conditions and a non-compensatory maturity index; empirical methods are systematic and replicable (coding protocols in Appendix B) but rely on non-random public samples (documented incidents, vendor/tool descriptions, public deployment disclosures) without counterfactuals or quantitative causal identification. SampleThree public/web-scale studies using publicly available sources and curated corpora: (1) a failure taxonomy mapping of documented production agent failures (authors report 82% of recorded failures are multi-layer trajectories — sample size not stated in the abstract), (2) a tooling capability analysis of 22 ecosystem tools (none offered full Layer 2 emergence coverage), and (3) a maturity scoring exercise of 35 public deployments (all scored in the lowest maturity band). Data derive from incident reports, vendor documentation, open-source repositories, public deployment writeups, and press disclosures; coding and labeling protocols are documented in Appendix B. Themesgovernance org_design adoption human_ai_collab GeneralizabilityFindings based on publicly documented incidents and disclosures are subject to publication and selection bias and may not represent the true universe of enterprise deployments or failures., Tooling and ecosystem coverage rapidly evolves; a 22-tool snapshot may be outdated quickly., Architecture examples and operational grounding drawn from Fortune‑50 scale platforms may not generalize to small/medium firms or public sector organizations., Regulatory and institutional constraints (e.g., EU AI Act) shape applicability — cross-jurisdiction differences limit generalizability., Empirical evidence is descriptive (taxonomy, coverage, maturity scoring) and does not establish causal impacts on productivity, employment, or firm performance.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The CASE framework assigns four governing disciplines to four scales of enterprise agentic AI: control theory to individual agents, complex adaptive systems theory to agent collectives, supervisory cybernetics to human-agent teams, and engineering operations/SRE to automated fleets. Governance And Regulation positive Fit between governance discipline and scale of agency
Reading fidelity high
Study strength medium
not reported
0.12
The paper reports that 82 percent of documented production agent failures were multi-layer trajectories rather than failures confined to a single governance layer. Ai Safety And Ethics positive Share of documented production agent failures involving multiple governance layers
Reading fidelity high
Study strength medium
82 percent
0.12
None of the 22 ecosystem tools evaluated by the paper offered full coverage of Layer 2, the emergence layer concerning agent collectives. Governance And Regulation negative Availability of full Layer 2 emergence-governance coverage in ecosystem tools
Reading fidelity high
Study strength medium
n=22
none of 22 tools
0.12
All 35 public deployments scored in the paper fell within the lowest CASE maturity band. Governance And Regulation negative CASE governance maturity level of public agentic AI deployments
Reading fidelity high
Study strength medium
n=35
all 35 deployments
0.12
The paper identifies an 'Emergence Gap': production risk is realized at the emergence layer while available tooling provides little capability there and organizational practice is largely absent. Ai Safety And Ethics negative Alignment between emergence-layer risk and governance capability
Reading fidelity high
Study strength medium
not reported
0.12
The paper argues that unaided human oversight of agentic AI fleets fails structurally because the variety of human responses is insufficient to match the behavioral variety generated by the agents. Governance And Regulation negative Adequacy of human oversight capacity relative to agent behavioral variety
Reading fidelity high
Study strength medium
not reported
0.12
The paper proposes that increasing fleet deployment automation can cause interaction variety to outpace oversight capacity, creating a zero-touch deployment paradox in which Layer 4 operational excellence strains Layers 2 and 3. Governance And Regulation negative Oversight capacity relative to interaction variety under increasing deployment automation
Reading fidelity high
Study strength speculative
not reported
0.02
The CASE framework extends the SRE error-budget concept beyond availability by incorporating decision quality and uses the remaining budget to modulate the permitted level of agent autonomy. Decision Quality positive Decision quality and permitted autonomy level as governed by an operational error budget
Reading fidelity high
Study strength speculative
not reported
0.02
The paper proposes a non-compensatory CASE maturity index in which the weakest governance layer receives majority weight, so stronger layers cannot compensate for a binding weakness. Organizational Efficiency positive Organizational governance maturity and identification of the binding capability layer
Reading fidelity high
Study strength speculative
not reported
0.02
The paper argues that certifying an agent artifact or model no longer certifies its runtime behavior because agents operate stochastically over open action spaces, adapt through context and memory, and can interact through shared state. Ai Safety And Ethics negative Validity of artifact-based certification as a proxy for runtime behavioral assurance
Reading fidelity high
Study strength medium
not reported
0.12

Notes