0 cumulative citations
View corpus contextRed-team benchmarks can reliably certify safety for common, high-frequency harms but are inadequate for rare catastrophic failures; a closed-form 'evidential ceiling' gives the exact harm-rate threshold where passing tests outweigh a single reproduced failure.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.
Summary
Main Finding
The paper formalizes what red-team (safety) evaluations can and cannot prove by defining an evidential ceiling: the maximum amount a single evaluation result can move posterior odds between an “acceptable-risk” hypothesis and an “elevated-risk” hypothesis given a fixed testing budget. It derives closed-form expressions for this ceiling for null (zero-harm) benchmark results, locates a precise boundary (in terms of harm rate, sample size, and decision threshold) above which feasible benchmarks can certify safety and below which they cannot, and shows that the key lever for evidence is discrimination between hypotheses (procedure-conditioned elicitation rates) rather than raw attack-success rates. Practical audits show current public benchmarks can certify common (high-frequency) harms but fall orders of magnitude short for rare, catastrophic harms.
Key Points
- Evidential ceiling: For any evidence-generating procedure E = (H, B, C, Λ), Ceil± is the supremum of log-likelihood-ratio magnitude feasible under the budget B and channel C. No rescoring or aggregation can increase this (data-processing inequality).
- Benchmark null result (n independent trials): likelihood ratio for zero harms is Λ0 = [(1 − pu)/(1 − ps)]^n where pu (H1) and ps = r·pu (H0) and r ∈ (0,1) is improvement ratio. In the rare-harm regime (small pu), Λ0 ≈ exp(−n·pu·(1 − r)), so null results become less informative as harms become rarer at fixed n.
- Two regimes:
- Above the boundary: if p > pmin(τ, Nmax, r) ≈ −ln τ / [Nmax(1 − r)], there exists n ≤ Nmax such that a zero-harm benchmark produces at least factor-τ update toward acceptable risk. Required n scales as O(1/p) and is computable beforehand.
- Below the boundary: if p < pmin, no benchmark with n ≤ Nmax and approximately independent Bernoulli trials can produce a null result meeting evidentiary threshold τ, regardless of prompt diversity.
- Crossing rate between evidence types: the rate p× where a clean sheet and a single observed harm carry equal evidence is p× ≈ ln(1/r) / [2 n (1 − r)]. Above p×, a clean sheet is stronger evidence; below p×, a single observed harm is stronger.
- Generalization to any elicitation procedure: replace pu, ps with q1 = P(elicit harm | H1) and q0 = P(elicit harm | H0). Evidence per trial is determined by discrimination κ = log2[(1 − q1)/(1 − q0)]. High q1 alone is uninformative if q0 is similarly high; attack-success rates (q1 alone) are thus the wrong metric.
- Decision-theoretic threshold τ: set by prior odds O0 and loss ratio L (cost of unsafe deployment relative to delay): τ = 1/(L·O0). This anchors required evidence to policy-relevant risk tolerance.
- Empirical audit summary (paper): assessed eight evaluation suites; found they are adequate for high-frequency harm categories but several orders of magnitude too small for rare/catastrophic categories.
Data & Methods
- Analytical, theoretical work: derivations and proofs produce closed-form formulas for Λ0, pmin, p×, and per-trial discrimination κ. Results are specializations of prior statistical bounds (e.g., zero-numerator bounds) but reinterpreted for evaluation design and decision thresholds.
- Main assumptions:
- Trials are approximately independent Bernoulli draws under the scoring rule and elicitation procedure being evaluated.
- Scoring rule, threat model, and elicitation procedure are fixed (so p or q are conditional on these choices).
- Budget constraint is an upper bound on independent trials Nmax (or on number of independent replications of the elicitation procedure).
- Adaptive campaigns are treated by redefining the trial unit to be the whole campaign (dependent probes inside a campaign do not count as independent trials).
- Audits: the paper evaluates (audits) eight public evaluation suites against the derived boundaries to assess practical adequacy. (Analytic results are illustrated with numeric examples: e.g., at n = 520, r = 0.5 the crossing rate p× ≈ 1.33×10−3; for Nmax = 1e5, τ = 0.5, r = 0.5, pmin ≈ 1.4×10−5.)
- Caveats and limits discussed:
- Distributional validity and structural generalization are separate necessary conditions for extrapolating test evidence to deployment (the paper formalizes Condition 1 — statistical sufficiency — and empirically discusses 2 and 3).
- Adaptive or model-mediated elicitation can change capacities but must be treated carefully: the trial unit becomes the full procedure, and model behavioral changes under detection can reduce discrimination.
- Measuring q1 requires a positive-control (reference model known to have the capability); without reporting q1/q0, null results cannot be meaningfully weighted.
Implications for AI Economics
- Cost scaling and infeasibility for rare harms:
- Required sample size scales as O(1/p). For rare, catastrophic harms (very small p), benchmark sample sizes and costs explode (often infeasible). Economically, trying to close safety gaps by linearly scaling benchmark size is an inefficient or impossible strategy for rare events.
- Efficient allocation: invest in procedure discrimination, not just more prompts:
- Improving elicitation discrimination (increasing q1 − q0 or κ) gives far more evidence per trial than increasing n. Adaptive, carefully designed elicitation that maximizes discrimination is a higher-return investment than naive benchmark growth.
- Reporting standards and market signals:
- Labs should report q1 and q0 (or κ), the trial unit, independence assumptions, and the decision-theoretic τ they target (derived from priors and loss ratios). Without this, benchmark clean sheets can be misinterpreted, producing misleading market/regulatory signals and perverse incentives to tune scoring rules.
- Regulatory and policy design:
- Regulators should require pre-specified hypotheses, target τ (linked to explicit priors and loss ratios), independent replication, and positive controls measuring q1. For harms below pmin, require non-benchmark evidence (system-level audits, operational monitoring, provable mitigations) rather than ever-larger benchmarks.
- Strategic behavior and incentives:
- Because the evidential value depends on scoring rules and elicitation procedures, producers may economically prefer to optimize scoring rules or construct reference models to make benchmarks look better cheaply. Policy should therefore mandate transparent metrics (q1, q0, Nmax) and independent auditing to avoid perverse optimizations.
- Cost–benefit trade-offs:
- Decision-makers can compute the marginal cost of raising evidence (increase n vs. raise κ) and compare against the societal cost L of a missed failure to choose economically efficient testing strategies. The paper’s closed-form expressions enable such cost-effectiveness calculations pre-deployment.
- Recommendation for economic actors:
- Shift resources from simply growing benchmark size toward (a) designing high-discrimination elicitation procedures, (b) building and reporting positive-control results, and (c) investing in non-benchmark evidence streams for rare catastrophic risks (monitoring, formal verification, operational constraints).
Summary takeaway: red-team benchmarks are informative but only for a computable set of propositions. Their evidentiary power is limited by sample-size economics (O(1/p)), by the discrimination of the elicitation procedure, and by decision-theoretic risk tolerances. For AI-economics decisions (investment, regulation, cost–benefit analysis), the paper supplies tractable formulas to quantify the costs of achieving a target evidentiary standard and shows that improving discrimination is often the more cost-effective lever than enlarging benchmark size.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget. Ai Safety And Ethics | mixed | evidential ceiling (largest multiplicative change in posterior belief caused by one result under fixed testing budget) |
Reading fidelity
high
Study strength
high
|
not reported
|
| We derive the evidential ceiling in closed form for the benchmark null result. Ai Safety And Ethics | mixed | closed-form expression for evidential ceiling under benchmark null result |
Reading fidelity
high
Study strength
high
|
not reported
|
| Above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Ai Safety And Ethics | positive | ability of a modest-sized benchmark to certify a category to a specified evidentiary standard (including comparative strength of a 'clean sheet' observation versus a single reproduced failure) |
Reading fidelity
high
Study strength
medium
|
modest size benchmark can certify above threshold harm rate (no numeric size reported in abstract)
|
| Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. Ai Safety And Ethics | negative | feasibility of passive benchmark to provide specified evidence of safety |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The crossing between the two regimes (above vs. below the harm-rate threshold) has a closed form. Ai Safety And Ethics | mixed | closed-form boundary (analytic expression) separating regimes where benchmarks can vs. cannot certify safety |
Reading fidelity
high
Study strength
high
|
not reported
|
| The bound is not specific to benchmarks: written in terms of a procedure's hypothesis-conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Ai Safety And Ethics | mixed | generality of bound (applies to adaptive/automated red teaming) and determinant of evidential worth (discrimination between hypotheses) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories. Ai Safety And Ethics | positive | adequacy of current benchmarks to provide specified evidential standard for high-frequency harm categories |
Reading fidelity
high
Study strength
medium
|
n=8
|
| Auditing eight evaluation suites against the boundary, we find that current benchmarks are several orders of magnitude short for rare, catastrophic harm categories. Ai Safety And Ethics | negative | magnitude of shortfall of current benchmarks relative to required evidence for rare catastrophic harms |
Reading fidelity
high
Study strength
medium
|
n=8
several orders of magnitude short
|
| Safety benchmarks are not uninformative: they are informative about a specific and computable set of propositions, and the discipline they need is to state which propositions they address. Ai Safety And Ethics | positive | informativeness of safety benchmarks about particular computable propositions (and requirement to state target propositions) |
Reading fidelity
high
Study strength
medium
|
not reported
|