The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new governance framework argues AI deployments should only make claims supported by explicit, contextual evidence and must record the institutional constraints that measurement cannot override; RISE AI ties system-level evaluation to who owns power, what is non-negotiable (e.g., dignity), and where claims must stop.

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good
Chawla, Nitesh V., Benanti, Paulo · September 10, 2026 · arXiv (Cornell University)
openalex commentary n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Chawla, Nitesh V. provider ID
  2. Benanti, Paulo provider ID

Semantic Scholar

Latest observation:

  1. Nitesh V. Chawla provider ID
  2. Paulo Benanti provider ID
The paper proposes RISE AI, an evidence architecture and 'rupture test' that bounds claims about Responsibility, Inclusivity, Safety, and Empowerment by linking system evaluation to preexisting institutional baselines and explicit measurement limits.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deployed, AI becomes an intervention in those conditions. It can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI must therefore evaluate both the system and the institutional rupture into which it is introduced. The move from principles to protocols is already underway. The EU AI Act, NIST AI RMF, ISO/IEC 42001, and assurance practices translate commitments into roles, requirements, records, oversight, and assessment. The harder questions are what these protocols actually establish, whose power they leave untouched, and where measurement must stop. Pope Leo XIV's Magnifica Humanitas provides a broader moral frame centered on dignity, technological power, and the common good. Drawing on that frame, we develop a rupture test that links institutional baselines to system evaluation. We distinguish evidence-bounded deployment, which limits claims to what has actually been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. Within those limits, RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment. Responsible AI requires better engineering, institutional repair, and continued moral and political judgment.

Summary

Main Finding

AI systems must be evaluated not only as technical artifacts but as interventions into pre‑existing institutional and relational failures. Responsible deployment requires (a) an explicit rupture test that links the institutional baseline to system evaluation; (b) evidence‑bounded claims about what a system actually accomplishes; and (c) measurement‑bounded governance that records normative constraints that favorable evidence cannot override. The paper operationalizes this with RISE — an evidence architecture for producing bounded claims about Responsibility, Inclusivity, Safety, and Empowerment — and argues that measurement and design cannot by themselves settle deeper political‑economic or dignity questions.

Key Points

  • AI reveals existing institutional weaknesses (responsiveness, belonging, accountability) and acts as an intervention that can repair, compound, substitute for, or conceal those weaknesses. Evaluation must address both revelation and intervention.
  • Rupture test: before deployment, identify the prior institutional/relational failure, the human/non‑AI baseline option, and whether the AI repairs/compounds/substitutes/conceals that failure.
  • From principles to protocols: the paper gives an operational chain — Principle → Design objective → System requirement → Implementation mechanism → Evaluation protocol → Bounded evidence claim.
  • Design patterns (connect governance objectives to engineering mechanisms):
    • Answerability‑by‑Design (role definitions, provenance, logs, redress)
    • Contestability‑by‑Design (usable appeals, protection from retaliation)
    • Agency‑Preserving Interfaces (useful friction; retain authority)
    • Context‑Aware Evaluation (population/workflow/version matching)
    • Evidence‑Bounded Deployment (claims limited to collected evidence)
  • Two fundamental limits of measurement:
    • Political economy/system boundaries: artifact‑level evaluation cannot by itself adjudicate ownership, financing, control, labor, or power asymmetries that shape legitimacy.
    • Dignity and construct validity: dignity is not a measurable latent variable; indicators can show violations of dignity but cannot measure dignity as such or trade it off against other outcomes.
  • Measurement‑bounded governance: alongside evidence‑based claims, deployments should record explicit constraints (uses, populations, relationships) that no benchmark result can license or override.
  • RISE architecture: produces versioned claim–evidence graphs linking bounded claims to indicators, evidence provenance, accountable owners, context, limitations, and expiry triggers; plus a context‑and‑power record (ownership, financing, labor, agenda‑setting) to make institutional order visible.

Data & Methods

  • Methodological approach: conceptual and normative analysis combining:
    • Policy and standards review (EU AI Act, NIST AI RMF, ISO/IEC 42001)
    • Measurement and validation theory (construct validity emphasis: validity applies to inference)
    • Design pattern synthesis from governance and engineering literatures
    • Assurance and documentation practices (assurance cases, provenance, logs)
    • Ethical and political framing using Pope Leo XIV’s Magnifica Humanitas as a moral/political lens on dignity, distribution of technological goods, and political economy
  • No primary empirical dataset: the paper is prescriptive/conceptual and constructs an operational architecture (RISE) for recording and assessing evidence and constraints.
  • Operational artifacts proposed: versioned claim–evidence graphs, context‑and‑power records, and protocol templates for accountability, contestation, and monitoring.

Implications for AI Economics

  • Make political economy explicit in economic assessments:
    • Evaluations must go beyond artifact performance to include ownership, control of data/compute, financing, labor inputs, and distribution of gains/losses. These institutional factors affect welfare, market structure, and bargaining power but are not captured by accuracy metrics.
  • Incorporate rupture tests into cost‑benefit and impact analyses:
    • When assessing deployment benefits, compare outcomes to the relevant non‑AI baseline (what alternative institutions or human services would have provided) and account for substitution/externalization effects.
  • Recognize limits of quantitative indicators:
    • Economists should not treat indicators (e.g., accuracy, uptake) as proxies for dignity, legitimacy, or social acceptability. Use indicators to support bounded causal inferences, but complement with normative, legal, and participatory inputs.
  • Measurement‑bounded governance affects incentives and markets:
    • Explicit constraints (prohibited uses, protected populations) change allowable product scope and can shift investment incentives. Regulators and firms will internalize compliance costs; assurance markets (auditing, evidence‑provision services) will grow in importance.
  • Valuation and procurement implications:
    • Public procurement and private investment should demand RISE‑style claim–evidence graphs and context‑and‑power records. This creates transaction costs but reduces asymmetric information about downstream social impacts.
  • Distributional and labor effects:
    • Analysts must account for labor displacement, retraining costs, and who pays for institutional repair (public vs. private), including environmental/resource costs (compute, energy) that affect social welfare.
  • Policy design and welfare modeling:
    • Welfare analyses should include non‑market values and institution‑level constraints (measurement‑bounded governance) as exogenous policy choices, not variables to be optimized purely by market outcomes.
  • Empirical agenda:
    • Develop standardized indicators for bounded claims (e.g., measurable contestation outcomes, override rates, recourse effectiveness) and methods for linking these to economic outcomes (employment, access, consumer surplus).
    • Study how context‑and‑power features (ownership concentration, procurement terms) mediate realized benefits and distributional outcomes of AI deployments.
  • Governance as an economic input:
    • Effective institutional capacity (e.g., trained reviewers, legal protections, monitoring) is itself an input with costs and returns; models should treat governance capacity as endogenous when predicting adoption and social outcomes.

Summary recommendation for economists: use the paper’s rupture test and RISE architecture to ground empirical and policy work in clearly specified baselines, bounded evidence claims, and explicit institutional constraints. This reduces the risk of conflating engineering performance with social legitimacy and helps align economic analysis with realistic, normatively informed assessments of AI’s social value.

Assessment

Paper Typecommentary Evidence Strengthn/a — This is a normative/framework paper with no empirical identification strategy or causal estimation; it proposes an evidence architecture (RISE) and conceptual tests rather than reporting empirical results. Methods Rigorn/a — The paper develops a conceptual framework and synthesizes policy, standards, and moral sources (EU AI Act, NIST, ISO, an encyclical) rather than applying empirical methods; arguments are coherent and grounded in existing documents but not subject to empirical validation here. SampleNo empirical sample; this is a conceptual/policy paper that draws on legal and standards texts (EU AI Act, NIST AI RMF, ISO/IEC 42001), assurance and governance literature, and the encyclical Magnifica Humanitas as a moral frame. Themesgovernance org_design GeneralizabilityNo empirical validation — applicability to real deployments is untested, Normative framing (use of a papal encyclical) may limit uptake across secular or diverse political contexts, Operationalization depends on institutional capacity and jurisdictional legal frameworks, Does not provide direct metrics for economic outcomes (productivity, wages), so linkage to economic effects is indirect, Implementation complexity means evidence claims may vary by sector, population, and system version

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI should be treated as both an intervention in, and a revelation of, pre-existing institutional or relational failures. Governance And Regulation mixed Whether AI deployments repair, compound, substitute for, or conceal institutional weaknesses.
Reading fidelity high
Study strength speculative
not reported
0.01
Responsible-AI evaluation should include a rupture test that identifies the institutional or relational failure preceding deployment, specifies a human or non-AI baseline, and determines whether the AI-mediated system repairs, compounds, substitutes for, or conceals that failure. Governance And Regulation positive Institutional repair or deterioration relative to a human or non-AI baseline.
Reading fidelity high
Study strength speculative
not reported
0.01
A general benchmark score can support only a narrow capability claim and cannot by itself establish responsibility, inclusivity, safety, or empowerment in a particular deployment context. Ai Safety And Ethics negative Transferability of benchmark evidence to context-specific responsible-AI outcomes.
Reading fidelity high
Study strength speculative
not reported
0.01
Technical and legal oversight mechanisms do not establish that human reviewers have the time, authority, understanding, and institutional protection needed for meaningful control. Governance And Regulation negative Meaningful human control over consequential AI-mediated decisions.
Reading fidelity high
Study strength speculative
not reported
0.01
Compliance with the EU AI Act alone does not establish empowerment, institutional repair, or a just distribution of technological power. Governance And Regulation negative Empowerment, institutional repair, and distributional justice after regulatory compliance.
Reading fidelity high
Study strength speculative
not reported
0.01
Evidence that a deployed system preserves user control, provides contestation, and communicates uncertainty does not establish that the broader distribution of technological power is just. Inequality negative Justice of the broader institutional and political-economic order served by an AI system.
Reading fidelity high
Study strength speculative
not reported
0.01
Dignity is not a measurable variable or indicator, and system-performance measures cannot establish, increase, or offset a person's inherent human worth. Ai Safety And Ethics null_result Whether system-performance indicators can measure or trade off human dignity.
Reading fidelity high
Study strength speculative
not reported
0.01
RISE AI does not measure dignity, empowerment, responsibility, inclusion, or safety as latent attributes; it evaluates whether specified evidence supports bounded claims about variable conditions. Governance And Regulation positive Validity and evidentiary support of bounded responsible-AI claims.
Reading fidelity high
Study strength speculative
not reported
0.01
Measurement-bounded governance should explicitly record constraints that remain binding regardless of favorable evidence, including prohibited uses, non-offsettable exclusions, and human relationships or capabilities that AI may not substitute away. Governance And Regulation positive Protection of normative, legal, and institutional boundaries against performance-based override.
Reading fidelity high
Study strength speculative
not reported
0.01
Responsible-AI claims should not exceed the evidence collected; evidence gaps, conflicts, and expiry conditions should trigger qualification, additional evaluation, remediation, or withdrawal when necessary. Governance And Regulation positive Scope and continued authorization of AI deployment claims.
Reading fidelity high
Study strength speculative
not reported
0.01
The appropriate unit of responsible-AI evaluation is the sociotechnical system rather than the AI model alone. Ai Safety And Ethics positive System-level responsibility and safety after deployment.
Reading fidelity high
Study strength speculative
not reported
0.01

Notes