The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Benchmarks that measure only correctness understate real-world risk: many AI errors are hard and costly for humans to detect, so evaluation should measure the verification effort required under realistic user budgets.

AI Evaluation Should Measure Verification Cost, Not Correctness Alone
Viviana Crescitelli, Generoso Immediato, Fabio Persia, Stefania Costantini · August 09, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Viviana Crescitelli unresolved corpus identity
  2. Generoso Immediato unresolved corpus identity
  3. Fabio Persia unresolved corpus identity
  4. Stefania Costantini unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Viviana Crescitelli provider ID
  2. Generoso Immediato provider ID
  3. Fabio Persia provider ID
  4. Stefania Costantini provider ID
The paper argues that model correctness is an incomplete reliability metric and proposes measuring verification cost relative to deployment verification budgets—defining Verification-Cost Errors (VCEs) as incorrect outputs that a sizable fraction of verifiers fail to detect within available verification resources.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those outputs. We argue that current evaluation metrics overlook a critical failure mode: Verification-Cost Errors (VCEs), defined as incorrect input-output pairs that a declared fraction of the verifier population fails to identify within the verification budget available in a given deployment context. Unlike standard notions of "hallucination", VCEs are defined operationally, by the failure of correct identification within budget rather than by any property of the output itself. Plausibility and authoritative presentation are hypothesised contributors to that failure, not defining conditions. To capture this asymmetry, we introduce the notion of verification cost relative to a deployment budget as an operational dimension that current evaluation does not routinely capture. The quantity is presented as a conceptual instrument rather than a finalized metric. Evidence from code generation and multi-modal document understanding shows that high benchmark accuracy can mask significant verification effort in practice. We therefore take the position that correctness alone is insufficient as a measure of reliability. AI evaluation should explicitly account for verification cost, reflecting whether errors can be detected under realistic resource constraints.

Summary

Main Finding

Current AI evaluation that reports correctness (benchmark accuracy) is insufficient. Reliability in deployment is determined not just by whether outputs are correct, but by how costly it is for users to detect incorrect outputs. The paper introduces Verification‑Cost Errors (VCEs): incorrect input–output pairs that a declared fraction of the verifier population fail to detect within the verification budget available in a deployment. Evaluation should therefore measure verification cost relative to deployment budgets, not correctness alone.

Key Points

  • Verification gap: Generating outputs (Cg) is often cheap; verifying them (Cv) can be orders of magnitude more expensive (Cv ≫ Cg). This asymmetry enables errors that are hard to detect in practice.
  • VCE definition: Operational — an incorrect output becomes a VCE when a specified fraction of verifiers cannot identify the error within the available verification budget.
  • Failure class emphasized: outputs that are stylistically fluent, locally coherent, and authority‑mimicking are particularly likely to produce VCEs (deceptive plausibility, not necessarily identified by traditional “hallucination” labels).
  • Limits of current framings:
    • Hallucination/factuality: Identifies incorrectness but not detectability under resource constraints.
    • Calibration/uncertainty: Helps triage but does not reduce verification cost and can miss systematic, high‑confidence errors.
    • Benchmarks: Treat all errors as uniform, hiding heterogeneity in detection effort.
    • Explanations and tool‑augmentation: May increase perceived trust or shift verification burden (e.g., verifying retrieved evidence or tool outputs), sometimes worsening human oversight cost.
  • Empirical motifs and examples:
    • Code generation: passes supplied tests but fails in edge cases; detecting requires extra testing or manual inspection.
    • OCR + LLM pipelines: character errors get smoothed into plausible but semantically wrong text, requiring source comparison.
    • RAG: grounding can create a veneer of authority; users must verify claim–source alignment.
    • Cited empirical findings (as motivating evidence): developer studies showing increased task time with AI assistance, RAG tools still hallucinating at nontrivial rates.
  • Proposal: verification‑aware benchmarking and explicit measures (authors propose constructs such as success‑probability and observed‑burden; treat verification cost as an axis of evaluation). The paper frames this as a conceptual instrument rather than a finished metric.

Data & Methods

  • Conceptual/formal constructs:
    • Define generation cost Cg(x,y) and verification cost Cv(x,y) for input x and output y.
    • Formalize VCEs as incorrect (x,y) pairs that fail detection by a declared fraction of a verifier population within a specified verification budget.
    • Introduce related measures (not fully standardized in the excerpt): success‑probability (probability a verifier detects the error within budget) and observed‑burden (empirical resources consumed to verify).
  • Empirical support (illustrative, not a large new dataset):
    • Case evidence from code generation, multi‑modal document understanding, and literature citations (e.g., Becker et al. 2025, Magesh et al. 2025, Dahl et al. 2024) documenting (i) hidden bugs or hallucinations that survive naive checks and (ii) increased human verification time in practice.
  • Methodological stance:
    • The paper is primarily conceptual and diagnostic; it proposes a benchmarking methodology in later sections (Section 6) to operationalize verification cost, but emphasizes the construct over a finalized metric.
  • Limitations noted by authors:
    • VCE measurement depends on the chosen verifier population and budget; creating standard, comparable protocols is nontrivial.
    • The work does not claim empirical prevalence numbers for VCEs across all domains; it identifies the measurement gap and proposes a way forward.

Implications for AI Economics

  • Hidden labor costs: Deploying generative AI shifts verification effort onto users or downstream workers. Economists and firms must account for verification labor as part of total cost of ownership (TCO) for AI systems.
  • Productivity accounting: Claims of productivity gains from AI may be overstated if they ignore increased verification time. Measured output per worker could decline once verification overheads are included.
  • Market for verification services and tools:
    • Demand for verification specialists, tooling, testing suites, provenance/audit systems, and certification services is likely to grow.
    • New business opportunities (and costs) arise for third‑party auditors and specialized verification platforms.
  • Pricing and procurement:
    • Procurement decisions should incorporate verification‑cost‑adjusted performance metrics (e.g., effective correct outcomes per verification hour).
    • Vendors may need to disclose expected verification costs or VCE rates for given deployment budgets; markets may evolve to price models accordingly.
  • Incentives and investment:
    • Model developers may optimize for benchmark accuracy rather than minimizing verification cost; incentive misalignment could persist unless benchmarks and procurement incorporate verification metrics.
    • Investors and managers evaluating ML projects should treat verification cost reduction as an explicit objective alongside accuracy and latency.
  • Liability and regulation:
    • Regulatory compliance and legal risk depend on whether errors are detectable within practicable budgets. Standards or certification regimes may require verification‑aware reporting.
    • Liability allocation (who pays for verification failures) becomes economically significant; insurers may price coverage differently for systems with high VCE risk.
  • Welfare and distributional effects:
    • Increased verification burdens may disproportionately affect smaller firms or under‑resourced users who lack time/expertise, creating asymmetric adoption or risk exposure.
    • If verification tasks concentrate in lower‑paid labor segments, there are distributional consequences even as model deployment proliferates.
  • Measurement and macro indicators:
    • Macro productivity statistics and returns to AI adoption should incorporate verification costs to avoid biased estimates of AI’s economic impact.
    • Benchmarking agencies and standard setters (public and private) could create verification‑aware metrics to guide policy and investment.

Practical takeaways for economists and decision‑makers: - Don’t take reported accuracy as a full measure of deployable reliability; ask how much time/expertise is needed to verify outputs. - Incorporate Cv (verification cost) into cost‑benefit analyses, procurement specifications, and ROI models. - Support or demand benchmarking standards that report verification burden or VCE rates under realistic verifier populations and budgets.

Limitations and open questions relevant to economic analysis: - Standardizing verifier populations, budgets, and tasks is required for comparability but may be contentious and domain‑specific. - Empirical quantification of VCE prevalence and the elasticity of verification cost with model scaling remain open research areas with important economic consequences.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is primarily conceptual and normative: it defines Verification-Cost Errors (VCEs), provides a formal framing and illustrative examples, and cites prior empirical work, but it does not present original causal identification or systematic empirical estimates of verification costs. Methods Rigormedium — The authors offer a clear formalization (Cg/Cv, operational definition of VCEs) and a coherent methodological proposal for verification-aware benchmarking, but they do not present implemented protocols, empirical measurements, or sensitivity analyses; arguments rely on plausible examples and prior literature rather than new empirical validation. SampleNo original empirical sample; the manuscript is conceptual and uses illustrative examples drawn from code generation, OCR→LLM document understanding, and RAG systems, and cites external empirical studies (e.g., Becker et al. 2025; Magesh et al. 2025) to motivate the framing. Themesproductivity human_ai_collab GeneralizabilityArgument is conceptual and may not map directly to all application domains without domain-specific operationalization of verification budget and verifier population., Verification cost depends on verifier expertise, tooling, and institutional processes; results will vary across firms, sectors, and regulatory regimes., No empirical calibration provided, so quantitative claims about prevalence or economic magnitude of VCEs are not generalizable., Proposed benchmarking methodology may be expensive or hard to standardize across benchmarks and languages/modalities.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI reliability cannot be measured by output correctness alone; evaluation should also account for the effort required to verify outputs under realistic resource constraints. Ai Safety And Ethics negative Reliability as a function of correctness and verification cost
Reading fidelity high
Study strength low
not reported
0.06
Verification-Cost Errors are incorrect input-output pairs that a declared fraction of verifiers fails to identify within the verification budget available in a deployment context. Ai Safety And Ethics negative Failure to detect incorrect AI outputs within a resource budget
Reading fidelity high
Study strength speculative
not reported
0.02
High benchmark accuracy can conceal substantial real-world verification effort, particularly in code generation and multi-modal document understanding. Organizational Efficiency negative Human effort required to verify AI-generated code and document outputs
Reading fidelity high
Study strength low
not reported
0.06
Proprietary retrieval-augmented legal AI tools continued to generate hallucinations in 17% to 33% of cases despite being marketed as 'hallucination-free.' Output Quality negative Hallucination rate in legal AI outputs
Reading fidelity high
Study strength medium
17% to 33% of cases
0.12
Experienced open-source developers took 19% longer to complete real tasks when AI assistance was permitted, while estimating afterward that AI assistance had made them 20% faster. Task Completion Time mixed Actual task completion time and developers' perceived speed under AI assistance
Reading fidelity high
Study strength medium
19% longer actual completion time; estimated 20% faster
0.12
Retrieval grounding does not necessarily eliminate verification burden and may shift it from assessing statement truthfulness to checking whether citations or retrieved evidence actually support the generated claims. Organizational Efficiency mixed Human verification burden for RAG-generated claims and sources
Reading fidelity high
Study strength low
not reported
0.06
Calibration and uncertainty estimation can help prioritize verification effort but do not reduce the labor or threshold required to establish correctness, especially for high-stakes outputs. Organizational Efficiency mixed Verification labor and allocation of verification effort
Reading fidelity high
Study strength low
not reported
0.06
Explanations and chain-of-thought outputs may increase verification burden by creating a semblance of logic that encourages users to rely on them too heavily. Ai Safety And Ethics negative Human verification burden and overreliance on AI explanations
Reading fidelity high
Study strength low
not reported
0.06
The paper conjectures that current AI scaling trends may widen rather than narrow the gap between generation cost and verification cost. Organizational Efficiency negative Gap between AI generation cost and human verification cost
Reading fidelity high
Study strength speculative
not reported
0.02

Notes