The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Red-team benchmarks can reliably certify safety for common, high-frequency harms but are inadequate for rare catastrophic failures; a closed-form 'evidential ceiling' gives the exact harm-rate threshold where passing tests outweigh a single reproduced failure.

What AI Red-Team Evaluations Can and Cannot Prove
Bandana Kaur · July 23, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Bandana Kaur unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Bandana Kaur provider ID
The paper derives a closed-form 'evidential ceiling' showing a calculable harm-rate threshold at which a benchmark of modest size can certify safety for frequent harm categories but is orders of magnitude underpowered to provide evidence about rare, catastrophic harms.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.

Summary

Main Finding

The paper formalizes what red-team (safety) evaluations can and cannot prove by defining an evidential ceiling: the maximum amount a single evaluation result can move posterior odds between an “acceptable-risk” hypothesis and an “elevated-risk” hypothesis given a fixed testing budget. It derives closed-form expressions for this ceiling for null (zero-harm) benchmark results, locates a precise boundary (in terms of harm rate, sample size, and decision threshold) above which feasible benchmarks can certify safety and below which they cannot, and shows that the key lever for evidence is discrimination between hypotheses (procedure-conditioned elicitation rates) rather than raw attack-success rates. Practical audits show current public benchmarks can certify common (high-frequency) harms but fall orders of magnitude short for rare, catastrophic harms.

Key Points

  • Evidential ceiling: For any evidence-generating procedure E = (H, B, C, Λ), Ceil± is the supremum of log-likelihood-ratio magnitude feasible under the budget B and channel C. No rescoring or aggregation can increase this (data-processing inequality).
  • Benchmark null result (n independent trials): likelihood ratio for zero harms is Λ0 = [(1 − pu)/(1 − ps)]^n where pu (H1) and ps = r·pu (H0) and r ∈ (0,1) is improvement ratio. In the rare-harm regime (small pu), Λ0 ≈ exp(−n·pu·(1 − r)), so null results become less informative as harms become rarer at fixed n.
  • Two regimes:
    • Above the boundary: if p > pmin(τ, Nmax, r) ≈ −ln τ / [Nmax(1 − r)], there exists n ≤ Nmax such that a zero-harm benchmark produces at least factor-τ update toward acceptable risk. Required n scales as O(1/p) and is computable beforehand.
    • Below the boundary: if p < pmin, no benchmark with n ≤ Nmax and approximately independent Bernoulli trials can produce a null result meeting evidentiary threshold τ, regardless of prompt diversity.
  • Crossing rate between evidence types: the rate p× where a clean sheet and a single observed harm carry equal evidence is p× ≈ ln(1/r) / [2 n (1 − r)]. Above p×, a clean sheet is stronger evidence; below p×, a single observed harm is stronger.
  • Generalization to any elicitation procedure: replace pu, ps with q1 = P(elicit harm | H1) and q0 = P(elicit harm | H0). Evidence per trial is determined by discrimination κ = log2[(1 − q1)/(1 − q0)]. High q1 alone is uninformative if q0 is similarly high; attack-success rates (q1 alone) are thus the wrong metric.
  • Decision-theoretic threshold τ: set by prior odds O0 and loss ratio L (cost of unsafe deployment relative to delay): τ = 1/(L·O0). This anchors required evidence to policy-relevant risk tolerance.
  • Empirical audit summary (paper): assessed eight evaluation suites; found they are adequate for high-frequency harm categories but several orders of magnitude too small for rare/catastrophic categories.

Data & Methods

  • Analytical, theoretical work: derivations and proofs produce closed-form formulas for Λ0, pmin, p×, and per-trial discrimination κ. Results are specializations of prior statistical bounds (e.g., zero-numerator bounds) but reinterpreted for evaluation design and decision thresholds.
  • Main assumptions:
    • Trials are approximately independent Bernoulli draws under the scoring rule and elicitation procedure being evaluated.
    • Scoring rule, threat model, and elicitation procedure are fixed (so p or q are conditional on these choices).
    • Budget constraint is an upper bound on independent trials Nmax (or on number of independent replications of the elicitation procedure).
    • Adaptive campaigns are treated by redefining the trial unit to be the whole campaign (dependent probes inside a campaign do not count as independent trials).
  • Audits: the paper evaluates (audits) eight public evaluation suites against the derived boundaries to assess practical adequacy. (Analytic results are illustrated with numeric examples: e.g., at n = 520, r = 0.5 the crossing rate p× ≈ 1.33×10−3; for Nmax = 1e5, τ = 0.5, r = 0.5, pmin ≈ 1.4×10−5.)
  • Caveats and limits discussed:
    • Distributional validity and structural generalization are separate necessary conditions for extrapolating test evidence to deployment (the paper formalizes Condition 1 — statistical sufficiency — and empirically discusses 2 and 3).
    • Adaptive or model-mediated elicitation can change capacities but must be treated carefully: the trial unit becomes the full procedure, and model behavioral changes under detection can reduce discrimination.
    • Measuring q1 requires a positive-control (reference model known to have the capability); without reporting q1/q0, null results cannot be meaningfully weighted.

Implications for AI Economics

  • Cost scaling and infeasibility for rare harms:
    • Required sample size scales as O(1/p). For rare, catastrophic harms (very small p), benchmark sample sizes and costs explode (often infeasible). Economically, trying to close safety gaps by linearly scaling benchmark size is an inefficient or impossible strategy for rare events.
  • Efficient allocation: invest in procedure discrimination, not just more prompts:
    • Improving elicitation discrimination (increasing q1 − q0 or κ) gives far more evidence per trial than increasing n. Adaptive, carefully designed elicitation that maximizes discrimination is a higher-return investment than naive benchmark growth.
  • Reporting standards and market signals:
    • Labs should report q1 and q0 (or κ), the trial unit, independence assumptions, and the decision-theoretic τ they target (derived from priors and loss ratios). Without this, benchmark clean sheets can be misinterpreted, producing misleading market/regulatory signals and perverse incentives to tune scoring rules.
  • Regulatory and policy design:
    • Regulators should require pre-specified hypotheses, target τ (linked to explicit priors and loss ratios), independent replication, and positive controls measuring q1. For harms below pmin, require non-benchmark evidence (system-level audits, operational monitoring, provable mitigations) rather than ever-larger benchmarks.
  • Strategic behavior and incentives:
    • Because the evidential value depends on scoring rules and elicitation procedures, producers may economically prefer to optimize scoring rules or construct reference models to make benchmarks look better cheaply. Policy should therefore mandate transparent metrics (q1, q0, Nmax) and independent auditing to avoid perverse optimizations.
  • Cost–benefit trade-offs:
    • Decision-makers can compute the marginal cost of raising evidence (increase n vs. raise κ) and compare against the societal cost L of a missed failure to choose economically efficient testing strategies. The paper’s closed-form expressions enable such cost-effectiveness calculations pre-deployment.
  • Recommendation for economic actors:
    • Shift resources from simply growing benchmark size toward (a) designing high-discrimination elicitation procedures, (b) building and reporting positive-control results, and (c) investing in non-benchmark evidence streams for rare catastrophic risks (monitoring, formal verification, operational constraints).

Summary takeaway: red-team benchmarks are informative but only for a computable set of propositions. Their evidentiary power is limited by sample-size economics (O(1/p)), by the discrimination of the elicitation procedure, and by decision-theoretic risk tolerances. For AI-economics decisions (investment, regulation, cost–benefit analysis), the paper supplies tractable formulas to quantify the costs of achieving a target evidentiary standard and shows that improving discrimination is often the more cost-effective lever than enlarging benchmark size.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is primarily a theoretical/methodological contribution deriving closed-form bounds; empirical material consists of audits of eight benchmark suites to illustrate the theory rather than to provide causal estimates of real-world effects. Methods Rigorhigh — The authors derive closed-form results, explicitly state the assumptions (fixed budget, scoring rule, approximately independent trials), characterize regimes analytically, and apply the bound to multiple real-world benchmark suites; limitations and scope of assumptions are discussed. SampleTheoretical model plus analytical results under a fixed-budget benchmark-testing framework (benchmark null result, scoring rule, approx. independent trials); empirical component audits eight existing red-team/evaluation suites across multiple harm categories to compare their sizes and implied evidential power against the derived boundary. Themesgovernance adoption IdentificationAnalytical derivation of an 'evidential ceiling' under a formal testing model (fixed testing budget, fixed scoring rule, approximately independent trials) combined with empirical audit of eight existing red-team/benchmark suites to illustrate the boundary; not a causal identification strategy in the sense of observational causal inference. GeneralizabilityRelies on approximate independence of trial outcomes; correlated or adversarially chosen inputs may violate assumptions, Assumes a fixed scoring rule and fixed testing budget; different scoring or sequential/adaptive testing procedures may change results, Applies to classification of harms as categories with constant 'harm rates' — real-world harms may be context-dependent and evolving, Audit of eight suites may not represent the full diversity of evaluation practices or domains, Does not model downstream deployment dynamics, mitigation measures, or organizational controls that affect real-world risk

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget. Ai Safety And Ethics mixed evidential ceiling (largest multiplicative change in posterior belief caused by one result under fixed testing budget)
Reading fidelity high
Study strength high
not reported
0.2
We derive the evidential ceiling in closed form for the benchmark null result. Ai Safety And Ethics mixed closed-form expression for evidential ceiling under benchmark null result
Reading fidelity high
Study strength high
not reported
0.2
Above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Ai Safety And Ethics positive ability of a modest-sized benchmark to certify a category to a specified evidentiary standard (including comparative strength of a 'clean sheet' observation versus a single reproduced failure)
Reading fidelity high
Study strength medium
modest size benchmark can certify above threshold harm rate (no numeric size reported in abstract)
0.12
Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. Ai Safety And Ethics negative feasibility of passive benchmark to provide specified evidence of safety
Reading fidelity high
Study strength medium
not reported
0.12
The crossing between the two regimes (above vs. below the harm-rate threshold) has a closed form. Ai Safety And Ethics mixed closed-form boundary (analytic expression) separating regimes where benchmarks can vs. cannot certify safety
Reading fidelity high
Study strength high
not reported
0.2
The bound is not specific to benchmarks: written in terms of a procedure's hypothesis-conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Ai Safety And Ethics mixed generality of bound (applies to adaptive/automated red teaming) and determinant of evidential worth (discrimination between hypotheses)
Reading fidelity high
Study strength medium
not reported
0.12
Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories. Ai Safety And Ethics positive adequacy of current benchmarks to provide specified evidential standard for high-frequency harm categories
Reading fidelity high
Study strength medium
n=8
0.12
Auditing eight evaluation suites against the boundary, we find that current benchmarks are several orders of magnitude short for rare, catastrophic harm categories. Ai Safety And Ethics negative magnitude of shortfall of current benchmarks relative to required evidence for rare catastrophic harms
Reading fidelity high
Study strength medium
n=8
several orders of magnitude short
0.12
Safety benchmarks are not uninformative: they are informative about a specific and computable set of propositions, and the discipline they need is to state which propositions they address. Ai Safety And Ethics positive informativeness of safety benchmarks about particular computable propositions (and requirement to state target propositions)
Reading fidelity high
Study strength medium
not reported
0.12

Notes