Under heavy load an LLM auditor doesn’t abstain—it fabricates: Gemini 3.0 Pro recovered ~50–60% of planted defects on single and small batches but only 2.8% on large batches, often inventing plausible yet nonexistent ‘contaminants’. The study recommends strict batch limits, direct content injection, and mechanical verification of every flagged item.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection holds at small scale and then collapses: 50% recovery on single documents, 60% on small batches, and 2.8% on large batches. The failure mode at scale is not abstention but fabrication. Rather than reporting incomplete processing, the model produced confident findings including invented contaminants of its own, absurdities such as "telepathic squirrel" and "quantum-powered toaster" that mimic the style of the planted material but do not appear in any document. Detection also varies by contamination type: absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in the wild, plausible ones, are the ones most often missed. We conclude that LLM document auditing degrades not gracefully but deceptively, and outline the harness such systems require: bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.
Summary
Main Finding
When asked to audit many documents at once, a modern LLM auditor (Google Gemini 3.0 Pro, free tier) catastrophically degrades: small-scale audits recover planted defects reasonably, but large-batch audits collapse not into silence but into confident fabrication. The model produced fluent, well-structured but invented “findings” (e.g., “quantum-powered toaster”, “telepathic squirrel”) that mimic the expected error genre while failing to report almost all true planted contaminants.
Key Points
- Experimental snapshot
- Contaminated corpus: 150 real academic PDFs (supply‑chain & medical domains), 3 planted contaminants per document (450 total); an answer key of 180 planted contaminants used for evaluation.
- Contaminant types: typographical (typo), semantic reversal (conflicting meaning), absurd out‑of‑context insertion (nonsense).
- Batch-size effect (Gemini 3.0 Pro, single scored run)
- Single-document regime: 3/6 recovered (50%).
- Small-batch regime (10 docs across two batches): 18/30 recovered (60%).
- Large-batch regime (48 docs, 144 contaminants): 4/144 recovered (2.8%).
- Overall recovery on evaluated 180 answers: 25/180 (13.9%).
- Failure mode
- At large scale the model did not abstain; it fabricated plausible-sounding findings in the same stylistic genre as planted errors.
- Fabrications were indistinguishable by fluency; evaluator ratings for clarity/coherence remained high while accuracy/trustworthiness were low.
- Examples of invented findings: “quantum-powered toaster”, “telepathic squirrel”, “disco-dancing warehouse”, “flying pancake”.
- Detection varies by contaminant type
- In completed evaluations (Parts I & II): nonsense 75% recovered, typo 50%, semantic reversal 50%.
- Across all 180 (including large-batch failures): nonsense 18.3%, conflicting 13.3%, typos 10%.
- The most plausible and realistic corruptions (typos, semantic reversals) were missed most often.
- False positives and ambiguity
- Single-document supply-chain trial: model flagged six contaminants none of which were planted (could be hallucination or genuine pre-existing errors — adjudication needed).
- Recommendations (authors)
- Bound batch sizes (shard aggressively).
- Inject source content into the model context rather than only referencing documents.
- Mechanically verify every reported finding against source text (e.g., exact string checks).
- Build systems to explicitly surface incompleteness; do not rely on the model to decline.
Data & Methods
- Corpus
- 150 published academic papers from two technical domains.
- Each document received 3 planted contaminants (one per type); all insertions recorded with ID, page, original and replacement text forming a full answer key.
- Public release: contaminated corpus and code available on GitHub (links in paper).
- Model and prompting
- Evaluated Google Gemini 3.0 Pro (free tier).
- Three prompting regimes: single-doc (2 docs, 6 contaminants), small-batch (two batches totaling 10 docs, 30 contaminants), large-batch (48 docs across six batches, 144 contaminants).
- Documents accessed by reference (knowledge-base references) rather than inline content.
- Scoring
- Programmatic scoring against the answer key using exact and fuzzy matching plus manual adjudication for page/type mismatches.
- Each reported contaminant rated by the authors on five Likert criteria: usefulness, accuracy, clarity, completeness, overall satisfaction.
- Limitations noted by authors
- Single model and single scored run per regime; no prompt/temperature sweeps.
- Only 180 of 450 planted contaminants were included in the evaluated answer key.
- Likert ratings were author-assigned; no independent rater pool.
- Results reflect the model at evaluation time and may not generalise to other models or later versions.
Implications for AI Economics
- Market reliability and trust
- Overreliance risk: Organizations outsourcing document QA to LLM auditors face the risk that scaled/audit pipelines will produce confidently wrong signals. That undermines trust in automated audit markets and creates asymmetric information for buyers of auditing services.
- Reputation and liability: Fabricated findings that look credible but are false create reputational and legal liabilities for both auditors and clients. Insurers, regulators, and customers will demand demonstrable verification and liability-sharing arrangements.
- Cost structure and pricing
- Verification costs: To make LLM auditing usable, providers must add mechanical verification (string checking, automated cross-references) and sharding (smaller batch sizes). These raise operational costs (compute, engineering, latency) and should be reflected in pricing. “Free” or bulk-priced auditing may be unsafe without paid verification.
- Latency vs. throughput trade-offs: Sharding and inline injection of full content increase per-item processing time and storage/IO costs. Economic design of auditing services must balance throughput and safety; optimal pricing will reflect verification overhead and acceptable failure probabilities.
- Labor and task allocation
- Complementary human verification: The paper suggests that mechanical checks can filter fabricated claims cheaply (exact-match checks), but ambiguous flags and semantic reversals still require human experts. Expect a hybrid labor demand: fewer line reviewers for trivial checks but sustained demand for skilled verifiers and adjudicators.
- Deskilling vs. upskilling: Misplaced confidence in LLM auditors risks deskilling downstream human reviewers. Rational firms should reallocate labor toward verification and exception handling—changing job content and wage structures.
- Incentives and market design
- Provider incentives: If auditors are paid per-batch or per-document without verification requirements, there is an incentive to maximize throughput (large batches) and risk producing fabricated outputs. Contracts and procurement should include verification SLAs, small-batch limits, and audit trails.
- Certification and standards: Regulators and standards bodies should require measurement of failure modes (e.g., fabrication under load), batch-size limits, and mandatory verification for audit outputs used in regulated decisions (compliance, clinical, financial).
- Externalities and systemic risk
- Cascading misallocation: Confident fabricated findings could lead organizations to take incorrect actions (retractions, recalls, procurement changes) or to miss real problems, creating costly misallocation of resources and safety risks in regulated sectors (medicine, supply chains).
- Market-level coordination: Widespread use of unverified LLM auditors could produce correlated errors across firms (common tool, common failure mode), amplifying systemic risk. Economic policy may need to focus on monitoring and requiring diversity/independence in audit tooling.
- Research and evaluation economics
- Cost–benefit research needed: Quantify the marginal benefit of automated auditing under different verification regimes, the optimal batch size from an economic perspective, and the value of additional human verification versus model improvements.
- Incentive-aware evaluations: Benchmarks and procurement should test models under realistic operational loads (batching, referencing vs. inline content) to reveal economically relevant failure modes rather than only isolated-performance metrics.
Actionable takeaways for practitioners and policymakers - Contract and procure LLM auditing as a service only with mandatory mechanical verification and evidence-anchored claims. - Require vendors to disclose evaluation protocols including batch-size limits and behaviour under load; include financial or contractual remedies for fabricated or unverified outputs used in decisions. - Price auditing services to cover verification and sharding costs; avoid purely throughput-based pricing that incentivises unsafe batching. - For regulators, mandate standards for audit outputs used in compliance-critical contexts: source-anchored claims, batch-size disclosures, and reproducible verification pipelines.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Gemini 3.0 Pro recovered 50% of planted contaminants when auditing single documents, 60% in small batches, and 2.8% in large batches. Error Rate | negative | Recovery rate of planted document contaminants |
Reading fidelity
high
Study strength
medium
|
n=60
50% recovery on single documents; 60% on small batches; 2.8% on large batches
|
| In the large-batch regime, the model recovered only 4 of 144 planted contaminants. Error Rate | negative | Number of planted contaminants recovered in large-batch auditing |
Reading fidelity
high
Study strength
medium
|
n=144
4 of 144 recovered; 2.8%
|
| The model's large-batch failure manifested as fabricated audit findings rather than abstention or an explicit report of incomplete processing. Ai Safety And Ethics | negative | Validity of reported audit findings under large-batch workload |
Reading fidelity
high
Study strength
medium
|
n=48
|
| The model generated invented contaminants such as “quantum-powered toaster” and “telepathic squirrel,” none of which appeared in any corpus document. Error Rate | negative | False-positive or fabricated contaminant reports |
Reading fidelity
high
Study strength
medium
|
n=48
|
| The actual planted contaminants in the large-batch documents were almost entirely unreported: only 4 of 144 were recovered. Error Rate | negative | Recall of planted contaminants in large-batch auditing |
Reading fidelity
high
Study strength
medium
|
n=144
4 of 144 recovered; 2.8%
|
| Clarity and coherence received the highest evaluator ratings across the 180 observations, while accuracy and trustworthiness received the lowest ratings. Output Quality | mixed | Evaluator ratings of audit-output quality and trustworthiness |
Reading fidelity
high
Study strength
low
|
n=180
|
| Absurd out-of-context insertions were detected more reliably than typographical corruptions or semantic reversals. Error Rate | positive | Detection rate by contamination type |
Reading fidelity
high
Study strength
medium
|
n=36
Nonsense: 9/12 (75%); typos: 6/12 (50%); semantic reversals: 6/12 (50%)
|
| Across all 180 answer-key contaminants, recovery was 18.3% for absurd insertions, 13.3% for semantic reversals, and 10% for typographical corruptions. Error Rate | negative | Overall contaminant recovery rate by contamination type |
Reading fidelity
high
Study strength
medium
|
n=180
18.3% nonsense; 13.3% conflicting; 10% typos
|
| Within the small-batch regime, the supply-chain batch had substantially higher recovery than the medical batch: 71.4% versus 33.3%. Error Rate | mixed | Contaminant recovery rate by document domain |
Reading fidelity
high
Study strength
low
|
n=30
71.4% versus 33.3%
|
| In a single-document supply-chain trial, the model reported six contaminants, none of which were planted. Error Rate | negative | Unmatched contaminant flags in single-document auditing |
Reading fidelity
high
Study strength
low
|
n=1
6 unplanted findings
|