0 cumulative citations
View corpus contextState-of-the-art vision–language models now out-see young adults at spotting AI-generated portraits—gpt-5.6-sol reached 92.8% balanced accuracy—yet machines are unstable to few-shot examples and carry extreme response biases, flipping about one in four verdicts when examples change.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.
Summary
Main Finding
Frontier vision–language models (VLMs) in July 2026 surpassed young adults at discriminating identity-matched AI-generated face portraits from real photos in sensitivity and balanced accuracy, but they do not match humans in calibration or decision balance. Top models (gpt-5.6-sol and claude-fable-5) now see more than young adults (d′ up to ≈3.4 vs ≈2.4) yet exhibit extreme and non-human response biases and substantial example-driven instability (≈1 in 4 item verdicts flip across few-shot example sets).
Key Points
- Benchmark outcome
- June-2026 cohort (14 models) matched but did not exceed the human young-adult ceiling (20s–30s). Best June single-pass ~87.1% balanced accuracy vs human mean 85.18% (young adults ≈88.5%).
- July-2026 cohort (5 new models) broke the ceiling: gpt-5.6-sol single-pass balanced accuracy 92.8% (five-draw mean 92.1%); claude-fable-5 averaged 91.9% and detected every AI image it saw (132/132).
- Sensitivity (signal-detection d′): July leaders gpt-5.6-sol d′ ≈ 3.13 and claude-fable-5 d′ ≈ 3.41 vs young-adult participant-mean ≈2.33 (pooled ≈2.41).
- Calibration and bias
- Humans across ages are nearly unbiased (criterion c ≈ 0). Models spread widely in bias: c ranged roughly −1.10 to +1.45 (AI-leaning to REAL-leaning).
- Examples: claude-opus-4.8 heavily over-called AI (bias +64.4 pp); llama-4-maverick heavily over-called REAL (bias −80.3 pp). New leaders also biased (fable-5 c ≈ −0.97; sol c ≈ +0.44).
- Balanced scoring masks such biases; a high-sensitivity but biased detector can still hurt downstream decisions.
- Stability and robustness
- Few-shot example sensitivity: models judged the same 198 test images six times (baseline + five different 4-example prefixes). Across the 19 models, mean per-item flip rate = 25.7% (≈1 in 4 items changed verdict at least once).
- Some models showed large within-model spread across draws (up to ~18 pp); Claude models were comparatively more stable.
- Confidence and metacognition
- June models tended to be overconfident and showed poor resolution. July leaders improved calibration and resolution (Brier scores: claude-fable-5 0.070, gpt-5.6-sol 0.073 vs June best 0.098).
- Even so, high confidence still accompanied many errors and model-stated rationales tended to align with the verdict rather than truth.
- Other observations
- Model performance did not neatly correlate with scale, price, or recency: some flagship models underperformed mid-tier siblings; improvements occurred rapidly (noticeable differences within ~4 weeks).
- The test task was narrow and controlled (identity-matched portraits, balanced class mix); performance in this setup does not automatically extend to open-world detection.
Data & Methods
- Stimuli
- 210 portrait images built from 70 FFHQ real identities and for each identity two AI-generated counterparts (ChatGPT-4o native generations and Imagen 3 via Gemini 2.5), mirrored to preserve identity. Main benchmark used 198 test images + 4 practice identities per session (practice excluded from scoring).
- Human reference
- Re-used data from an earlier human study (n = 1,667 adults aged 20–69). Human sessions: 4 practice trials with feedback, then 20 REAL/AI trials (10 real, 5 ChatGPT-4o, 5 Imagen 3). Reported human mean balanced accuracy 85.18%; young adults in their 20s ≈88.5%.
- Model roster & protocol
- 19 VLMs from 9 providers (OpenAI, Anthropic, Google, xAI, Alibaba, Zhipu, Mistral, Meta, Moonshot) were evaluated via OpenRouter at temperature 0.
- Each model got the same 4 labelled practice images as a few-shot prefix (2 real, 1 ChatGPT-4o, 1 Imagen 3), then judged each test image once, returning: REAL/AI verdict, 0–100 confidence, and a one-sentence rationale.
- Models were run on a baseline draw and five additional draws that rotated which 4 practice examples were used to measure example-sensitivity and stability.
- Scoring & analyses
- Main metric: balanced accuracy weighted to human session composition (0.5·real + 0.25·ChatGPT-4o + 0.25·Imagen 3).
- Bootstrapped 20-trial resampling reproduced the human-session composition for comparable confidence intervals.
- Signal-detection measures (d′ sensitivity, criterion c) computed with log-linear correction; calibration measured by mean confidence vs accuracy and Brier score; flip rates measured across runs; rationale content and diagnosticity analyzed (3,738 one-sentence rationales collected).
- Controls to separate example-effect vs run-to-run noise were included.
Implications for AI Economics
- Market and productization
- Detection automation is now economically feasible for narrower, controlled forensic tasks: top VLMs can outperform typical young-adult human reviewers in sensitivity, which supports deployment in platform content-moderation pipelines, verification services, and forensic tools.
- Rapid frontier improvement (weeks) implies short product development cycles and frequent model refreshes; firms must plan for continuous integration, benchmarking, and monitoring costs.
- Risk allocation and externalities
- Model bias (tendency to over-call AI or REAL) creates asymmetric economic harms:
- False positives (flagging real profiles/ads) impose user friction, reputational cost, and potential legal exposure for platforms — substantial for advertisers and creators.
- False negatives (missing AI fakes) propagate misinformation and fraud, with social and financial costs.
- Balanced-accuracy metrics hide such asymmetries; procurement and regulation should require reporting class-specific errors and calibrated thresholds tailored to application costs.
- Reliability, trust, and liability
- Example-sensitivity and ~25% flip rates across practice prompts signal brittleness. For high-stakes uses (law enforcement, ID verification, financial KYC), model instability undermines reliability and increases audit/insurance costs.
- Firms may demand models with tunable decision criteria and documented calibration, or hybrid human-in-the-loop designs to manage residual errors; this affects labor demand for reviewers (shift from bulk screening to exception handling) and staffing cost structures.
- Competition, pricing, and product differentiation
- Performance did not strictly follow price or scale; vendors can differentiate on stability, calibration, and tunability rather than raw accuracy alone. Markets may emerge for “calibrated detectors” and audit services.
- Rapid frontier gains create winner-take-most dynamics for short windows, but persistent brittleness keeps room for competing products emphasizing robustness and explainability.
- Policy and measurement
- Economic models of platform risk and social cost should incorporate not only detection accuracy but calibration, bias, and stability — metrics that determine real-world cost trade-offs (blocking legitimate content vs letting fake content pass).
- Standardized, human-comparable benchmarks (as used here) are valuable for procurement, regulation, and liability assessments. Regulators and certifiers should require public reporting of bias (per-class error), calibration (Brier score), and stability (flip rate under example perturbations).
- Research and investment priorities
- Investment in methods for post-hoc calibration, tunable thresholds, robust few-shot prompting, and domain-adaptive detectors will likely yield high ROI for deployment.
- Insurance and legal frameworks need to account for algorithmic bias and instability; economic models of adoption should weigh the reduced marginal need for human reviewers against increased governance, auditing, and reputational risk costs.
Overall, the paper shows the economics of detection tools is shifting from “can models reach human accuracy?” to “how reliably and safely can they be deployed given their bias, calibration, and brittleness?” Decision-makers should evaluate detectors on a multidimensional frontier (sensitivity, bias, calibration, stability) and design incentives, procurement rules, and insurance to reflect those dimensions.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Among the July-2026 model releases, gpt-5.6-sol achieved 92.8% balanced accuracy, exceeding the mean accuracy of adults in their 20s (88.5%). Decision Quality | positive | Balanced accuracy in distinguishing real portraits from AI-generated portraits |
Reading fidelity
high
Study strength
high
|
n=198
92.8% balanced accuracy versus 88.5%
|
| gpt-5.6-sol and claude-fable-5 exceeded the young-adult human sensitivity level for distinguishing real from AI-generated portraits. Decision Quality | positive | Signal-detection sensitivity (d′) for separating real and AI-generated portraits |
Reading fidelity
high
Study strength
high
|
n=198
d′ = 3.13 for gpt-5.6-sol and d′ = 3.41 for claude-fable-5 versus young-adult d′ ≈ 2.41
|
| The July model leaders did not match human calibration: claude-fable-5 and gpt-5.6-sol exhibited response biases, whereas human observers remained near an unbiased criterion. Ai Safety And Ethics | negative | Response criterion and balance between REAL and AI judgments |
Reading fidelity
high
Study strength
high
|
n=198
c = −0.97 for claude-fable-5 and c = +0.44 for gpt-5.6-sol versus humans near c = 0
|
| Changing the four labelled few-shot practice images caused approximately one quarter of model per-item judgments to change. Ai Safety And Ethics | negative | Stability of binary portrait-classification judgments across few-shot example sets |
Reading fidelity
high
Study strength
medium
|
n=19
25.7% fleet-wide flip rate; 24.7% strict REAL↔AI reversals
|
| The July leaders were more accurately calibrated than the June models according to the difference between stated confidence and balanced accuracy. Ai Safety And Ethics | positive | Confidence calibration, defined as the gap between stated confidence and observed balanced accuracy |
Reading fidelity
high
Study strength
medium
|
n=198
gpt-5.6-sol calibration gap of +0.1 percentage points versus June-model gaps up to +29 percentage points
|
| The two July leaders had the best reported Brier scores among the evaluated models, indicating improved probabilistic prediction quality relative to the June cohort. Ai Safety And Ethics | positive | Brier score for confidence-based predictions |
Reading fidelity
high
Study strength
medium
|
n=198
Brier score 0.070 for claude-fable-5 and 0.073 for gpt-5.6-sol versus June best 0.098
|
| Model performance varied substantially across the 19 evaluated VLMs, ranging from 58.3% to 92.8% balanced accuracy. Decision Quality | mixed | Balanced accuracy across vision-language models |
Reading fidelity
high
Study strength
high
|
n=19
58.3% to 92.8% balanced accuracy
|
| Human accuracy in the underlying portrait-detection study declined with age, from approximately 88% among adults in their 20s to approximately 66% among adults in their 60s. Decision Quality | negative | Accuracy in detecting AI-generated face portraits |
Reading fidelity
high
Study strength
medium
|
n=1667
Approximately 88% in the 20s versus approximately 66% in the 60s
|