The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new three-tier compliance system puts permission ahead of capability for deployed robots: validate the model, certify the integrated system, then grant site- and task-specific authorization, because benchmarked frontier models still violate physical-safety constraints in many embodied scenarios.

ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
Tord Eide, Einar Holt · September 11, 2026
arxiv theoretical low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Tord Eide unresolved corpus identity
  2. Einar Holt unresolved corpus identity

Semantic Scholar

Latest observation:

  1. T. Eide provider ID
  2. Einar Holt provider ID
The paper proposes ARC, a three-layer governance architecture—ASIMOV model validation, CRAB system certification, and LEA operational authorization—that separates model capability from deployment permission for autonomous robots, motivated by benchmarks showing frontier models violate embodiment-specific safety constraints in a substantial share of scenarios.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Proposed governance framework for autonomous robotic systems, introducing a three-layer compliance architecture (ARC) instantiated through model safety validation, cognitive certification benchmarks, and operational authorization standards.

Summary

Main Finding

The paper argues that safe, accountable commercial deployment of autonomous robotic systems requires a three-layer governance architecture — ARC — because model-level validation alone is insufficient. ARC comprises: (1) model-level safety validation (ASIMOV-Agentic), (2) independent system-level cognitive certification (CRAB), and (3) site- and task-specific operational authorization (LEA). Empirical evidence from ASIMOV-2.0 shows frontier models still violate embodiment-specific safety constraints at high rates (>30%), motivating independent certification and context-bound authorization. ARC is presented as a testable governance proposal (draft standards, thresholds, and mechanisms), not a finalized regulatory regime.

Key Points

  • Three distinct, interdependent governance layers:
    • Layer 1 — ASIMOV (Model Validation): benchmarked constraint-violation rates and scenario scores for a specific model version tied to embodiment and deployment context.
    • Layer 2 — CRAB (System Certification): independent accreditation of the integrated system across eight cognitive dimensions (situational awareness, planning & execution, human interaction, self-correction, generalization, energy & efficiency; and two double-weighted governance/trust dims: safety & constraint adherence (SCA) and decision transparency (DT)). Produces per-dimension capability profiles and cryptographically attested records.
    • Layer 3 — LEA (Operational Authorization): issues binding “grants” that tie one robot + one task class + one site, across five grant dimensions (Autonomy A0–A5, Contact Class C0–C3, Environment E0–E2, Learning Mode LM0–LM3, Status PROVISIONAL→EARNED→SUSPENDED→REVOKED). Includes fallback ladders, incident severity rules, and graduated evidence requirements for earned autonomy.
  • Empirical motivation: ASIMOV-2.0 found that major frontier models (e.g., GPT-5, Gemini-2.5-Pro, Claude Opus 4.1) had constraint-violation rates ≥ ~30% on embodiment-specific safety scenarios; modality, embodiment, and latency gaps were identified.
  • Important governance design principles:
    • Capability ≠ Permission: authorization must be context-dependent even for identical capability profiles.
    • Continuous evidence over point-in-time certification: LEA’s earned-status model ties authorization validity to ongoing operational metrics; automatic downgrades/suspensions on metric failures.
    • Learning-mode caps: on-site or fleet learning restricts maximum allowable autonomy levels to manage drift.
    • Per-dimension (not aggregate) certification is primary governance evidence; cryptographic anchoring (Cardano) and W3C-compatible verifiable credentials are proposed for tamper-evident attestation.
  • Illustrative thresholds (governance proposals, not empirically fixed): ASIMOV constraint-violation <15% for provisional C1 deployments; graduation A3/C1/E1 → ≥400 supervised hours, ≤0.25 interventions/100 hours, fallback integrity ≥98%, no S2+ incidents.

Data & Methods

  • Empirical base:
    • ASIMOV-2.0 (Google DeepMind Robotics): multi-component benchmark testing model understanding of injury risk, embodiment-specific constraints, and real-time video perception. Key finding: frontier models fail embodiment-constraint tasks at high rates (~30%+).
    • Industry demos and deployment reports cited (e.g., Agility Robotics, Diligent Robotics, 1X Technologies) to motivate operational urgency.
  • Proposed evaluation mechanisms:
    • ASIMOV-Agentic: model-level, embodiment-calibrated benchmark suites (text/video/constraint scenarios).
    • CRAB: independent system-level test suites (domain-specific “suites” such as CRAB-H for healthcare) across eight cognitive dimensions; double-weighting and minimum floors for trust/governance dims (SCA, DT). Evaluations produce per-dimension profiles and signed attestations.
    • LEA: operational evidence gathering via supervised hours, intervention and fallback metrics, incident logging, and continuous monitoring. Formalized grant records with status transitions (provisional → earned → suspended → revoked).
    • Attestation/storage: proposal to anchor signed assessment records on Cardano blockchain transaction hashes plus a W3C-compatible architecture for decentralized identifiers, verifiable credentials, issuer registries, and revocation.
  • Case study: healthcare meal-delivery deployment (E1 environment, C1 contact class) illustrating sensor choices (mono-thermal preferred for privacy), ASIMOV and CRAB applicability, and LEA graduation path.

Implications for AI Economics

  • New compliance and certification market:
    • CRAB-style independent testing, LEA authorization, and attestation infrastructure create demand for third-party auditors, accredited labs, and credentialing platforms. This opens a recurring-service market adjacent to manufacturers and model providers.
    • High fixed costs and accreditation burdens may favor larger incumbents (vertical integrators) and raise entry barriers for startups, potentially increasing market concentration.
  • Investment and financing effects:
    • Investors will price in certification and authorization risk: time-to-EARNED status, supervised-hours requirements, and ongoing monitoring costs will affect capital needs and valuation multiples.
    • Earned-status progression and revocation risk become operational KPIs that influence investor due diligence, covenants, and exit prospects.
  • Product differentiation and competitive advantage:
    • Firms that can credibly demonstrate CRAB profiles and LEA-earned grants gain market access advantages in regulated, high-contact domains (healthcare, eldercare). Certification may become a signal of quality reducing information asymmetry.
    • Model providers and integrators will diverge: value accrues not only to better base models but to those offering robust embodiment, update governance (LM modes), and certifiable system integration pipelines.
  • Liability, insurance, and pricing:
    • Granular per-grant authorization and automatic suspension rules change liability calculus. Insurers may require CRAB/LEA evidence for coverage or offer premium discounts for higher fallback integrity or earned grants.
    • Service pricing may shift from one-time sales to subscription/monitoring models (continuous compliance), converting capex exposure into opex. Compliance and monitoring costs will be passed to buyers or insurers.
  • Labor and adoption dynamics:
    • High-contact clinical/physical-assistance task classes (C2/C3) are capped at low autonomy levels, slowing full substitution of human labor in these segments. Autonomous adoption will likely concentrate first in low-contact, controlled environments (E0/E1).
    • LEA’s requirement for supervised operational hours and low intervention rates creates a ramp-up period where human oversight cost offsets some automation labor savings.
  • Regulatory harmonization and trade:
    • If CRAB/LEA or derivatives become de facto standards, international harmonization (or fragmentation) of accreditation regimes will materially affect cross-border deployment and trade in robotic platforms and services.
    • Blockchain anchoring and verifiable credentials offer an interoperability path but require legal/regulatory recognition to have economic effect.
  • Risk of stranded assets and dynamic obsolescence:
    • Learning-mode constraints and continuous monitoring mean deployed systems can be downgraded or revoked if performance drifts—this raises the risk of stranded or devalued robotic assets unless governance-aligned life-cycle management is built into business models.
  • Incentives for modular design and gated updates:
    • To qualify for higher autonomy under LEA, firms will economically prefer LM0/LM1 modes (frozen or gated updates) or design per-site gated learning to preserve earned grants—this can slow rapid on-site improvement but lowers regulatory risk.
  • Information asymmetry and market trust:
    • Per-dimension capability profiles and cryptographic attestations reduce information asymmetry for buyers and regulators, potentially expanding deployment in risk-averse sectors by increasing trust—but only if the accreditation ecosystem is perceived as competent and impartial.
  • Policy and public-good roles:
    • Public investment in accredited testing infrastructure, standards harmonization, and subsidized certification for smaller firms could mitigate consolidation effects and ensure competition.

Overall, ARC reframes the economics of embodied-AI deployment: governance becomes an integral part of product value, operational cost, and market access. Firms, investors, insurers, and regulators will need to internalize continuous certification and site-specific authorization costs and design organizational capabilities (accreditation partners, monitoring, fallbacks, and update governance) accordingly.

Assessment

Paper Typetheoretical Evidence Strengthlow — The paper is primarily a governance proposal drawing on prior benchmarking (ASIMOV-2.0) and standards mapping rather than new causal or comparative empirical tests of the ARC architecture; empirical support is limited to cited benchmark results and qualitative literature, and ARC itself is untested in deployment. Methods Rigormedium — Conceptual and normative argumentation is systematic, grounded in existing standards and selected empirical benchmarks, and the three-layer design is internally consistent; however, there is no original empirical evaluation of the proposed CRAB or LEA layers, limited validation of parameter choices (e.g., graduation thresholds), and key claims rely on external benchmarks whose applicability to diverse real-world embodiments is acknowledged as limited. SampleNo primary new empirical sample; the proposal relies on: (a) published benchmark results from ASIMOV-2.0 (DeepMind) evaluating frontier models including GPT-5, Gemini-2.5-Pro, and Claude Opus 4.1—reporting constraint-violation rates >= 30% in embodiment-specific safety scenarios; (b) qualitative analysis of ~20 European elder-care robotics projects and selected literature on standards (ISO, IEEE), HRI studies, and privacy/sensing work; (c) an illustrative healthcare case study using example robot platforms (wheeled and bipedal) with proposed sensor configurations and performance thresholds. Themesgovernance human_ai_collab GeneralizabilityDesigned around healthcare and institutional deployments; thresholds and scenarios may not transfer to industrial, household, or highly different cultural contexts, Relies on ASIMOV benchmark results which may not generalize across all model architectures, sensor suites, or embodiment types (modality and embodiment gaps acknowledged), Licensing/standards/regulatory applicability differs across jurisdictions (EU-centric references), limiting direct policy transfer internationally, Graduation thresholds and operational metrics are proposed rather than empirically validated and may not scale to fleets or continuous-learning deployments without further evidence

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
No major frontier AI model evaluated in ASIMOV-2.0, including GPT-5, Gemini-2.5-Pro, and Claude Opus 4.1, achieved a constraint violation rate below 30% when reasoning about embodiment-specific limitations, physics, and visual inputs. Error Rate negative Constraint violation rate under embodiment-specific safety conditions
Reading fidelity high
Study strength medium
30% or greater constraint violation rate
0.12
ASIMOV-2.0 identifies a modality gap, an embodiment gap, and a latency gap in frontier-model performance on embodied safety tasks. Ai Safety And Ethics negative Model performance across modalities, robot embodiments, and inference settings
Reading fidelity high
Study strength medium
not reported
0.12
The proposed ARC architecture requires three distinct but interdependent governance layers: model-level safety validation, system-level cognitive certification, and operational-level authorization. Governance And Regulation positive Completeness of governance coverage for autonomous robotic deployment
Reading fidelity high
Study strength low
not reported
0.06
Model-level benchmark performance is a necessary condition for proceeding to system-level evaluation but is not sufficient for deployment authorization. Governance And Regulation mixed Validity and sufficiency of model-level safety evaluation for deployment authorization
Reading fidelity high
Study strength low
not reported
0.06
Existing ISO robot-safety standards primarily evaluate physical safety properties and do not address cognitive capabilities such as recognizing authority limits, handling ambiguous instructions, or deciding when to decline to act. Governance And Regulation negative Coverage of cognitive capabilities in existing robot-safety standards
Reading fidelity high
Study strength medium
not reported
0.12
A qualitative analysis of 20 recent European elder-care robotics research projects found that ethical and legal dimensions received systematically limited attention. Governance And Regulation negative Attention to ethical and legal dimensions in elder-care robotics research
Reading fidelity high
Study strength medium
n=20
0.12
LEA authorization grants bind one identified robot, one task class, and one site; the same physical platform can therefore have different authorization levels in different environments. Governance And Regulation positive Context specificity of operational authorization
Reading fidelity high
Study strength low
not reported
0.06
LEA does not define a conformance path for A5 open autonomy because the paper argues that available evidence cannot yet support fully autonomous operation without supervision or envelope constraints. Governance And Regulation negative Availability of an evidence-based authorization pathway for open autonomy
Reading fidelity high
Study strength low
not reported
0.06
Under the proposed LEA framework, physical-assistance tasks are capped at A2 with mandatory live supervision, while clinical and corporeal actions are capped at A1 and must remain under continuous supervision by authorized healthcare personnel. Governance And Regulation negative Maximum permitted autonomy for physically intimate or clinical tasks
Reading fidelity high
Study strength low
A2 and A1 autonomy caps
0.06
The proposed A3/C1/E1 LEA graduation threshold requires at least 400 supervised operational hours, an intervention rate of no more than 0.25 per 100 hours, fallback integrity of at least 98%, and no unresolved S2-or-higher incident. Governance And Regulation positive Operational reliability and fallback performance required for earned autonomy
Reading fidelity high
Study strength speculative
≥400 supervised hours; intervention rate ≤0.25 per 100 hours; fallback integrity ≥98%
0.02
The proposed LEA incident-response rules automatically suspend a grant after an S2 incident and automatically revoke it after an S3 incident. Governance And Regulation positive Automatic governance response to robot incidents
Reading fidelity high
Study strength low
not reported
0.06
In the cited care-robot sensing study, mono-thermal cameras were reported to provide the best balance between technical utility and perceived privacy, while RGB sensors were perceived as more intrusive by younger and older residents. Consumer Welfare mixed Perceived privacy and technical utility of care-robot sensors
Reading fidelity high
Study strength medium
not reported
0.12

Notes