The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

State-of-the-art vision–language models now out-see young adults at spotting AI-generated portraits—gpt-5.6-sol reached 92.8% balanced accuracy—yet machines are unstable to few-shot examples and carry extreme response biases, flipping about one in four verdicts when examples change.

Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration
Sunwhi Kim, Sunyul Kim, Meounggun Jo, Jini Tae · August 31, 2026
arxiv descriptive high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Sunwhi Kim unresolved corpus identity
  2. Sunyul Kim unresolved corpus identity
  3. Meounggun Jo unresolved corpus identity
  4. Jini Tae unresolved corpus identity
In a like-for-like benchmark, the latest vision–language models (notably gpt-5.6-sol and claude-fable-5) now exceed young adults at detecting AI-generated face portraits, but remain unstable across few-shot examples and poorly calibrated compared with human judges.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.

Summary

Main Finding

Frontier vision–language models (VLMs) in July 2026 surpassed young adults at discriminating identity-matched AI-generated face portraits from real photos in sensitivity and balanced accuracy, but they do not match humans in calibration or decision balance. Top models (gpt-5.6-sol and claude-fable-5) now see more than young adults (d′ up to ≈3.4 vs ≈2.4) yet exhibit extreme and non-human response biases and substantial example-driven instability (≈1 in 4 item verdicts flip across few-shot example sets).

Key Points

  • Benchmark outcome
    • June-2026 cohort (14 models) matched but did not exceed the human young-adult ceiling (20s–30s). Best June single-pass ~87.1% balanced accuracy vs human mean 85.18% (young adults ≈88.5%).
    • July-2026 cohort (5 new models) broke the ceiling: gpt-5.6-sol single-pass balanced accuracy 92.8% (five-draw mean 92.1%); claude-fable-5 averaged 91.9% and detected every AI image it saw (132/132).
    • Sensitivity (signal-detection d′): July leaders gpt-5.6-sol d′ ≈ 3.13 and claude-fable-5 d′ ≈ 3.41 vs young-adult participant-mean ≈2.33 (pooled ≈2.41).
  • Calibration and bias
    • Humans across ages are nearly unbiased (criterion c ≈ 0). Models spread widely in bias: c ranged roughly −1.10 to +1.45 (AI-leaning to REAL-leaning).
    • Examples: claude-opus-4.8 heavily over-called AI (bias +64.4 pp); llama-4-maverick heavily over-called REAL (bias −80.3 pp). New leaders also biased (fable-5 c ≈ −0.97; sol c ≈ +0.44).
    • Balanced scoring masks such biases; a high-sensitivity but biased detector can still hurt downstream decisions.
  • Stability and robustness
    • Few-shot example sensitivity: models judged the same 198 test images six times (baseline + five different 4-example prefixes). Across the 19 models, mean per-item flip rate = 25.7% (≈1 in 4 items changed verdict at least once).
    • Some models showed large within-model spread across draws (up to ~18 pp); Claude models were comparatively more stable.
  • Confidence and metacognition
    • June models tended to be overconfident and showed poor resolution. July leaders improved calibration and resolution (Brier scores: claude-fable-5 0.070, gpt-5.6-sol 0.073 vs June best 0.098).
    • Even so, high confidence still accompanied many errors and model-stated rationales tended to align with the verdict rather than truth.
  • Other observations
    • Model performance did not neatly correlate with scale, price, or recency: some flagship models underperformed mid-tier siblings; improvements occurred rapidly (noticeable differences within ~4 weeks).
    • The test task was narrow and controlled (identity-matched portraits, balanced class mix); performance in this setup does not automatically extend to open-world detection.

Data & Methods

  • Stimuli
    • 210 portrait images built from 70 FFHQ real identities and for each identity two AI-generated counterparts (ChatGPT-4o native generations and Imagen 3 via Gemini 2.5), mirrored to preserve identity. Main benchmark used 198 test images + 4 practice identities per session (practice excluded from scoring).
  • Human reference
    • Re-used data from an earlier human study (n = 1,667 adults aged 20–69). Human sessions: 4 practice trials with feedback, then 20 REAL/AI trials (10 real, 5 ChatGPT-4o, 5 Imagen 3). Reported human mean balanced accuracy 85.18%; young adults in their 20s ≈88.5%.
  • Model roster & protocol
    • 19 VLMs from 9 providers (OpenAI, Anthropic, Google, xAI, Alibaba, Zhipu, Mistral, Meta, Moonshot) were evaluated via OpenRouter at temperature 0.
    • Each model got the same 4 labelled practice images as a few-shot prefix (2 real, 1 ChatGPT-4o, 1 Imagen 3), then judged each test image once, returning: REAL/AI verdict, 0–100 confidence, and a one-sentence rationale.
    • Models were run on a baseline draw and five additional draws that rotated which 4 practice examples were used to measure example-sensitivity and stability.
  • Scoring & analyses
    • Main metric: balanced accuracy weighted to human session composition (0.5·real + 0.25·ChatGPT-4o + 0.25·Imagen 3).
    • Bootstrapped 20-trial resampling reproduced the human-session composition for comparable confidence intervals.
    • Signal-detection measures (d′ sensitivity, criterion c) computed with log-linear correction; calibration measured by mean confidence vs accuracy and Brier score; flip rates measured across runs; rationale content and diagnosticity analyzed (3,738 one-sentence rationales collected).
    • Controls to separate example-effect vs run-to-run noise were included.

Implications for AI Economics

  • Market and productization
    • Detection automation is now economically feasible for narrower, controlled forensic tasks: top VLMs can outperform typical young-adult human reviewers in sensitivity, which supports deployment in platform content-moderation pipelines, verification services, and forensic tools.
    • Rapid frontier improvement (weeks) implies short product development cycles and frequent model refreshes; firms must plan for continuous integration, benchmarking, and monitoring costs.
  • Risk allocation and externalities
    • Model bias (tendency to over-call AI or REAL) creates asymmetric economic harms:
    • False positives (flagging real profiles/ads) impose user friction, reputational cost, and potential legal exposure for platforms — substantial for advertisers and creators.
    • False negatives (missing AI fakes) propagate misinformation and fraud, with social and financial costs.
    • Balanced-accuracy metrics hide such asymmetries; procurement and regulation should require reporting class-specific errors and calibrated thresholds tailored to application costs.
  • Reliability, trust, and liability
    • Example-sensitivity and ~25% flip rates across practice prompts signal brittleness. For high-stakes uses (law enforcement, ID verification, financial KYC), model instability undermines reliability and increases audit/insurance costs.
    • Firms may demand models with tunable decision criteria and documented calibration, or hybrid human-in-the-loop designs to manage residual errors; this affects labor demand for reviewers (shift from bulk screening to exception handling) and staffing cost structures.
  • Competition, pricing, and product differentiation
    • Performance did not strictly follow price or scale; vendors can differentiate on stability, calibration, and tunability rather than raw accuracy alone. Markets may emerge for “calibrated detectors” and audit services.
    • Rapid frontier gains create winner-take-most dynamics for short windows, but persistent brittleness keeps room for competing products emphasizing robustness and explainability.
  • Policy and measurement
    • Economic models of platform risk and social cost should incorporate not only detection accuracy but calibration, bias, and stability — metrics that determine real-world cost trade-offs (blocking legitimate content vs letting fake content pass).
    • Standardized, human-comparable benchmarks (as used here) are valuable for procurement, regulation, and liability assessments. Regulators and certifiers should require public reporting of bias (per-class error), calibration (Brier score), and stability (flip rate under example perturbations).
  • Research and investment priorities
    • Investment in methods for post-hoc calibration, tunable thresholds, robust few-shot prompting, and domain-adaptive detectors will likely yield high ROI for deployment.
    • Insurance and legal frameworks need to account for algorithmic bias and instability; economic models of adoption should weigh the reduced marginal need for human reviewers against increased governance, auditing, and reputational risk costs.

Overall, the paper shows the economics of detection tools is shifting from “can models reach human accuracy?” to “how reliably and safely can they be deployed given their bias, calibration, and brittleness?” Decision-makers should evaluate detectors on a multidimensional frontier (sensitivity, bias, calibration, stability) and design incentives, procurement rules, and insurance to reflect those dimensions.

Assessment

Paper Typedescriptive Evidence Strengthhigh — Carefully controlled, directly comparable human–model benchmark with a large human reference sample (n=1,667), 19 contemporary VLMs, repeated draws to measure instability, bootstrapped session scoring, and signal-detection and calibration analyses; limitations are mostly about scope of stimuli and generators rather than internal validity. Methods Rigorhigh — The experiment matches human and model protocols (same stimuli, balanced trial mix, few-shot examples), uses temperature 0 for repeatability, repeats draws to measure example sensitivity, applies bootstrapping and cluster corrections, and reports multiple metrics (balanced accuracy, d′/criterion, calibration, flip rates); limitations include a modest stimulus set (198 test images), only two generator families for the AI images, and single-session judgments per draw. SampleStimuli: 210 face portraits (70 real photos from FFHQ and 140 identity-matched AI counterparts generated in July 2025 by ChatGPT-4o and Imagen 3), with 198 main test images per model after excluding practice identities. Human comparison: previously collected dataset of 1,667 adults aged ~20–69 who completed 20-trial sessions. Models: 19 vision–language models from nine providers (14 collected June 2026, 5 added July 2026) each given the same 4 labelled few-shot examples then asked to classify each test image once at temperature 0, returning REAL/AI, 0–100 confidence, and a one-sentence rationale; analyses included five varied few-shot draws, bootstrapped 20-trial session resampling, and signal-detection metrics. Themesgovernance human_ai_collab GeneralizabilityStimulus domain limited to frontal face portraits (FFHQ) at 512×512 px; findings may not generalize to other image types (scenes, screenshots, heavily post-processed photos) or lower-quality social-media imagery., AI-generated images come from two generators (ChatGPT-4o and Imagen 3, July 2025); results may not extend to other generative models, newer generators, or post-processing pipelines., Lab-style single-image binary judgments with balanced class mix are not the same as real-world detection (streaming feeds, contextual cues, multiple images per identity)., Models were queried via OpenRouter at temperature 0 and with a specific few-shot protocol; different access methods, prompts, or temperatures could change behavior., Human comparison uses a single prior human dataset (demographics/representativeness beyond age not fully described), which may limit cross-population generalization., Findings are a snapshot of rapidly evolving models and may be out-of-date as model architectures, training data, and safety layers change.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Among the July-2026 model releases, gpt-5.6-sol achieved 92.8% balanced accuracy, exceeding the mean accuracy of adults in their 20s (88.5%). Decision Quality positive Balanced accuracy in distinguishing real portraits from AI-generated portraits
Reading fidelity high
Study strength high
n=198
92.8% balanced accuracy versus 88.5%
0.3
gpt-5.6-sol and claude-fable-5 exceeded the young-adult human sensitivity level for distinguishing real from AI-generated portraits. Decision Quality positive Signal-detection sensitivity (d′) for separating real and AI-generated portraits
Reading fidelity high
Study strength high
n=198
d′ = 3.13 for gpt-5.6-sol and d′ = 3.41 for claude-fable-5 versus young-adult d′ ≈ 2.41
0.3
The July model leaders did not match human calibration: claude-fable-5 and gpt-5.6-sol exhibited response biases, whereas human observers remained near an unbiased criterion. Ai Safety And Ethics negative Response criterion and balance between REAL and AI judgments
Reading fidelity high
Study strength high
n=198
c = −0.97 for claude-fable-5 and c = +0.44 for gpt-5.6-sol versus humans near c = 0
0.3
Changing the four labelled few-shot practice images caused approximately one quarter of model per-item judgments to change. Ai Safety And Ethics negative Stability of binary portrait-classification judgments across few-shot example sets
Reading fidelity high
Study strength medium
n=19
25.7% fleet-wide flip rate; 24.7% strict REAL↔AI reversals
0.18
The July leaders were more accurately calibrated than the June models according to the difference between stated confidence and balanced accuracy. Ai Safety And Ethics positive Confidence calibration, defined as the gap between stated confidence and observed balanced accuracy
Reading fidelity high
Study strength medium
n=198
gpt-5.6-sol calibration gap of +0.1 percentage points versus June-model gaps up to +29 percentage points
0.18
The two July leaders had the best reported Brier scores among the evaluated models, indicating improved probabilistic prediction quality relative to the June cohort. Ai Safety And Ethics positive Brier score for confidence-based predictions
Reading fidelity high
Study strength medium
n=198
Brier score 0.070 for claude-fable-5 and 0.073 for gpt-5.6-sol versus June best 0.098
0.18
Model performance varied substantially across the 19 evaluated VLMs, ranging from 58.3% to 92.8% balanced accuracy. Decision Quality mixed Balanced accuracy across vision-language models
Reading fidelity high
Study strength high
n=19
58.3% to 92.8% balanced accuracy
0.3
Human accuracy in the underlying portrait-detection study declined with age, from approximately 88% among adults in their 20s to approximately 66% among adults in their 60s. Decision Quality negative Accuracy in detecting AI-generated face portraits
Reading fidelity high
Study strength medium
n=1667
Approximately 88% in the 20s versus approximately 66% in the 60s
0.18

Notes