The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Most inference-time controls already exist in production, but they are fragile: 15 of 20 mechanisms have commercial substrates today, yet none withstand high-capability or state-level adversaries and many enforcement primitives are defeated by fine-tuning, implying regulators cannot rely on training-compute thresholds alone.

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance
Samar Ansari · September 09, 2026
arxiv descriptive n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Samar Ansari unresolved corpus identity
The paper builds a feasibility taxonomy of 20 inference-time AI governance mechanisms, rates them using vendor evidence and an adversary matrix, finds 15 have commercial technical substrates today but none are adequate against high-capability state-level adversaries, and maps substitutions between inference- and hardware-stage instruments.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Compute governance today is a governance of training: the thresholds, reporting requirements, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. That picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. This paper asks which mechanisms are available once the regulatory object shifts from the training run to the inference call. We develop a feasibility taxonomy of twenty inference-time mechanisms across monitoring, verification, and enforcement, each rated on a four-point readiness scale against a documented four-vendor evidence base. We then stress the taxonomy against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and map each mechanism to four governance scenarios (domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance). Fifteen of the twenty mechanisms have commercial technical substrates in production today, although governance-grade assurance and adversarial robustness vary substantially. The adversary analysis shows that this readiness holds only against a cooperative deployer and a low-to-medium-capability user: no mechanism rates adequate against a high-capability state-level deployer, and fine-tuning removes the model-internal components of the enforcement cluster, although platform-external controls can persist. A substitution analysis connects the taxonomy to a companion hardware paper as a conditional substitution principle describing when inference-stage and hardware-stage mechanisms provide comparable regulatory coverage under stated conditions. A second-rater reliability check on a random subset of the readiness ratings returned a quadratic-weighted Cohen's kappa of 0.74.

Summary

Main Finding

Training-focused compute governance is necessary but no longer sufficient: a substantial portion of AI capability now emerges at inference time (via inference scaling, agentic scaffolding, and compression to consumer hardware). Ansari develops a structured feasibility taxonomy of 20 inference-time governance mechanisms (monitoring, verification, enforcement), rates them against real-world vendor evidence on a four-point readiness scale, and evaluates them against a two-dimensional adversary model and four governance scenarios. While 15 of 20 mechanisms have commercial technical substrates in production today, their assurance and adversarial robustness are uneven: mechanisms generally work against cooperative deployers and low-to-medium-capability users but fail against high-capability, state-level deployers, and many enforcement primitives are undermined by fine-tuning.

Key Points

  • Why inference matters

    • Capability can migrate to deployment through: (1) inference scaling (repeated sampling, longer chains of thought, ensembles), (2) agentic scaffolding (tool use, memory, external resources), and (3) compression to consumer hardware.
    • Training-compute thresholds (e.g., EU AI Act, US reporting regimes) are silent on inference and can be bypassed via inference-side techniques.
  • Taxonomy and readiness

    • 20 inference-time governance mechanisms organized into Monitoring, Verification, and Enforcement clusters.
    • New verification primitive added: V7 — chain-of-thought monitorability.
    • Four-point readiness scale: 1) currently deployable (commercial substrate in production), 2) near-term, 3) requires R&D, 4) speculative.
    • Conservative rating rule: where evidence conflicts, the more conservative rating is used.
  • Empirical readiness

    • 15 of 20 mechanisms have commercial technical substrates in production at major providers.
    • Readiness varies by adversary role and capability: most mechanisms are adequate only against cooperative deployers and low-to-medium-capability end users.
    • No mechanism in the taxonomy is rated adequate against a high-capability state-level deployer.
    • Fine-tuning and post-distribution modification can remove model-internal enforcement mechanisms; platform-external controls may still function.
  • Adversary model & scenarios

    • Two-dimensional adversary model: 3 capability tiers × 4 adversary roles (developer, deployer/integrator, end user, fine-tuner/scaffold-builder).
    • Mechanisms mapped to four governance scenarios: domestic regulation, bilateral/multilateral coordination (treaties, shared certs), industry self-regulation, and compute-marketplace/platform governance.
  • Substitution analysis

    • A convergence/substitution analysis links inference-stage mechanisms to hardware-stage mechanisms (from a companion hardware governance paper) and formulates a conditional substitution principle specifying when inference controls can substitute for hardware controls.
  • Method reliability & scope

    • Ratings derived from a documented four-vendor evidence base (commercial providers); consumer-hosted/open-weight deployments are treated as coverage limitations.
    • Second-rater check on a random subset (7 mechanisms) yielded quadratic-weighted Cohen’s kappa = 0.74 (substantial agreement); unweighted kappa = 0.53 (moderate).
    • Full adversary and scenario matrices are provided in appendices.

Data & Methods

  • Method type: feasibility taxonomy and qualitative-quantitative readiness assessment grounded in vendor documentation and academic literature.
  • Mechanism set: 20 inference-time mechanisms classified into Monitoring, Verification, Enforcement (detailed definitions and per-mechanism descriptions in Section 3).
  • Evidence base: documented vendor/provider sources from four major commercial providers (snapshots saved; listed in Appendix D).
  • Readiness rating procedure:
    • Four-point scale (1 = currently deployable; 4 = speculative).
    • Conservative decision rule: when sources disagree, adopt more conservative rating; when academic sources silent on adversarial conditions, rate against adversarial (not cooperative) case.
  • Adversary model: two-dimensional—capability tiers (low/medium/high) crossed with adversary role (developer, deployer, user, fine-tuner/scaffold-builder). Ratings are scoped by adversary surface.
  • Validation: independent second rater assessed a random subset of mechanisms (7/20); inter-rater reliability measured with quadratic-weighted Cohen’s kappa = 0.74. Disagreements documented in appendices.
  • Coverage limitations: focus on commercial-provider-mediated inference; consumer hardware and unfederated open-weight deployment are handled as cases where platform-mediated mechanisms lose force (flagged per mechanism).

Implications for AI Economics

  • Regulatory coverage and market leakage

    • Training-centric regulation can create regulatory arbitrage: firms (or states) can achieve frontier capabilities cheaply via inference-side techniques or local (consumer) deployment. This weakens the efficacy of ex ante capital controls on training compute and shifts the emphasis to post-training controls and marketplace governance.
    • Open-weight models and consumer-hardware deployment reduce chokepoints, increasing the likelihood of spillovers and making market-based enforcement (platform intermediaries, KYC for API purchasers) less effective. Economically, this favors decentralization and may reduce marginal enforcement costs but increases negative externalities.
  • Market structure and gatekeeper power

    • Platform- and marketplace-mediated inference controls (attestation, per-request monitoring, KYC) increase the regulatory leverage of major cloud providers and marketplaces, potentially reinforcing their market power. Policymakers will face trade-offs between relying on these intermediaries (effective enforcement) and preserving competition and innovation.
    • If regulators require platform-mediated controls to manage inference risk, new compliance costs and switching frictions could entrench incumbents or raise barriers to entry.
  • Compliance costs, incentives, and strategic behavior

    • Firms will weigh the cost of building inference governance (monitoring infrastructure, attestations, secure enclaves) against the expected compliance and liability costs. Where enforcement is weak against high-capability adversaries, firms may face incentives to underinvest in governance, particularly for deployments aimed at opaque or adversarial actors.
    • Fine-tuning capability undermines many model-internal enforcement mechanisms. Economic actors could intentionally delegate risky modifications to downstream fine-tuners to evade obligations—creating moral hazard across the developer–deployer–user chain.
  • International coordination and public-good problems

    • No single inference-time technical mechanism is adequate vs. high-capability state actors; effective mitigation requires international coordination (mutual recognition, treaty verification, export controls). This elevates collective-action and public-good problems: countries must invest in joint monitoring and enforcement to avoid a global race-to-the-bottom.
    • Economic sanctions, trade measures, or cross-border platform restrictions may become the practical tools to deter state-level misuse—each with significant geopolitical and economic costs.
  • Innovation, investment, and R&D priorities

    • The finding that 15/20 mechanisms have commercial substrates suggests near-term feasibility for many governance tools, but the assurance and adversarial robustness gaps indicate R&D priorities: hardening verification (e.g., robust attestation, chain-of-thought monitorability), improving enforcement that survives fine-tuning, and privacy-preserving monitoring.
    • Markets (venture, corporate R&D) are likely to fund solutions where regulatory demand or liability risk creates commercial opportunity (e.g., attestation providers, secure inference enclaves, certified toolchains). Public R&D may be needed for mechanisms that are socially valuable but not privately profitable (robust verification vs. state-level adversaries).
  • Liability, insurance, and pricing of inference

    • If inference governance is required, providers may price in monitoring/attestation services and KYC; risk-based pricing of API access (or tiered access) could emerge, with higher compliance costs for high-risk use cases.
    • Insurance markets may develop to underwrite inference-related harms, but insurers will demand measurable controls and auditability—driving standardization of verification mechanisms. The uneven robustness against powerful adversaries could make insurance coverage limited or expensive for high-risk deployments.
  • Policy design tradeoffs

    • Regulators must choose layers to regulate: training compute, inference calls, marketplaces, or hardware. The paper’s conditional substitution principle implies economic trade-offs: hardware-layer controls can be more durable against some forms of circumvention but are costly and less feasible globally; inference-layer controls are implementable via intermediaries but vulnerable to decentralization and fine-tuning.
    • Differential regulatory instruments (e.g., mandatory attestation + export restrictions + targeted sanctions) may be needed to cover different adversary/scenario combinations, but this creates compliance complexity and potential uneven regulatory burdens across firms and countries.

Overall, the paper signals that economic analysis of AI governance must expand beyond training-compute-focused models to incorporate inference-stage dynamics: altered incentive structures, new gatekeepers, arbitrage paths, and public-good and international coordination challenges.

Assessment

Paper Typedescriptive Evidence Strengthn/a — The paper is a conceptual/feasibility taxonomy rather than an empirical causal study; it compiles vendor documentation and literature to rate technical readiness but does not attempt causal identification or estimate effects. Methods Rigormedium — The author uses a clear, replicable taxonomy, an explicit four-point readiness scale, conservative rating rules, vendor-anchored evidence, and reports inter-rater reliability on a random subset (quadratic-weighted kappa = 0.74). However, the ratings rely heavily on vendor documentation and expert judgment, the second-rater sample is small (7 of 20 mechanisms), adversary assumptions are stylized, and there is no independent empirical validation or robustness checks beyond the limited re-rating. SampleA structured review and taxonomy of 20 inference-time governance mechanisms (grouped into monitoring, verification, enforcement), rated against a vendor evidence base drawn from four major commercial providers; literature synthesis; a random subset (7 mechanisms) independently re-rated for inter-rater reliability; vendor documentation snapshots and appendices provide supporting materials and filled adversary and scenario matrices. Themesgovernance adoption GeneralizabilityLimited to major commercial providers and platform-mediated deployments; consumer-hosted and unfederated open-weight deployments are treated as coverage gaps rather than directly analysed., Findings depend on rapidly evolving vendor implementations and dated documentation snapshots, so readiness ratings may change quickly., Legal, privacy, and jurisdictional constraints differ across countries and may change feasibility in practice., Adversary model is stylized (three capability tiers × four roles) and may not capture all real-world attacker behaviours or hybrid roles., No empirical validation of mechanism effectiveness in adversarial field conditions; results are conditional on authors' conservative assumptions.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The paper develops a feasibility taxonomy of 20 inference-time AI governance mechanisms organized into monitoring, verification, and enforcement categories. Governance And Regulation positive Coverage and feasibility of inference-time governance mechanisms
Reading fidelity high
Study strength medium
n=20
20 mechanisms
0.18
Fifteen of the 20 inference-time governance mechanisms have commercial technical substrates in production at major providers. Adoption Rate positive Commercial availability of inference-time governance mechanisms
Reading fidelity high
Study strength medium
n=20
Fifteen of the twenty mechanisms
0.18
The reported readiness of inference-time governance mechanisms applies only against a cooperative deployer and a low-to-medium-capability user; no mechanism is rated adequate against a high-capability state-level deployer. Governance And Regulation negative Governance-mechanism adequacy under adversarial conditions
Reading fidelity high
Study strength medium
n=20
No mechanism rated adequate against a high-capability state-level deployer
0.18
Fine-tuning removes the model-internal components of the enforcement cluster, although platform-external controls can remain effective. Governance And Regulation mixed Persistence of enforcement controls after fine-tuning
Reading fidelity high
Study strength medium
n=20
0.18
The second-rater reliability check produced a quadratic-weighted Cohen's kappa of 0.74. Governance And Regulation positive Inter-rater reliability of readiness ratings
Reading fidelity high
Study strength medium
n=7
quadratic-weighted Cohen's kappa = 0.74
0.18
The reliability estimate is sensitive to the weighting choice and small sample size: unweighted kappa was 0.53 rather than 0.74. Governance And Regulation mixed Robustness of inter-rater reliability estimates
Reading fidelity high
Study strength medium
n=7
unweighted kappa = 0.53
0.18
Existing EU, US, and UK compute-governance threshold regimes attach obligations to training compute but are silent on inference. Governance And Regulation negative Regulatory coverage of inference-time compute
Reading fidelity high
Study strength medium
n=3
3 of 3 regimes silent on inference
0.18
Training-compute thresholds fail to capture high-capability deployments that rely on above-optimal inference-time compute scaling. Governance And Regulation negative Regulatory coverage of capability generated through inference-time scaling
Reading fidelity high
Study strength medium
1026 versus 1024 training FLOPs in the illustrative example
0.18
Inference-time capability can arise from repeated sampling, longer reasoning, larger ensembles, agentic scaffolding, and model compression onto consumer hardware, weakening the assumption that training compute alone determines capability. Automation Exposure negative Relationship between training compute and deployed AI capability
Reading fidelity high
Study strength medium
not reported
0.18
Consumer-hardware and unfederated open-weight deployment limit the coverage of platform-mediated inference-governance mechanisms. Governance And Regulation negative Coverage of platform-mediated governance mechanisms across deployment settings
Reading fidelity high
Study strength medium
not reported
0.18
The paper's readiness ratings measure availability of a commercial technical substrate, not completeness of governance-grade assurance or adversarial robustness. Governance And Regulation mixed Interpretability of technology-readiness ratings
Reading fidelity high
Study strength high
n=20
0.3

Notes