The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Predictive saliency models miss much of what audiences actually look at: across a nationally quotaed sample viewing news photos, a simple centered map outperforms leading trained models, and model errors disproportionately affect older, Black, and politically extreme viewers, raising fairness risks for systems that use a single attention map.

Human versus Computer Vision
Elena Sirotkina · August 10, 2026
arxiv descriptive high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Elena Sirotkina unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Elena Sirotkina provider ID
On 83 circulating news photographs viewed by a quota-representative panel of 3,023 US adults (11.4M webcam gaze points), a blank centered Gaussian map predicts viewers' gaze better than six leading saliency models, and the residual predictive power of trained models systematically favors younger, White, and politically moderate viewers over older, Black, and ideologically extreme ones.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see everyone, and this study supplies the standard by which such a claim should be judged.

Summary

Main Finding

On circulating news photographs viewed in situ by a demographically representative US sample (webcam gaze), commercial and research saliency networks largely rely on a universal central bias. A blank centered Gaussian map (no training) predicts viewers’ gaze better than every trained network tested. The trained models add only a thin margin beyond that center, and that margin transfers unevenly across demographic groups — favoring younger, White, and politically moderate viewers and failing for older, Black, and ideologically extreme viewers. A small panel of real viewers of the same photo already matches or exceeds model performance for many use cases, suggesting practical remedies.

Key Points

  • Data scale and sample
    • 3,023 US adults recruited to national quotas (age, gender, education, race, income, partisanship).
    • 11.4 million webcam gaze samples across 61,458 viewings of 83 circulating Getty news photographs.
  • Models evaluated
    • Six predictors: four deep saliency networks (DeepGazeIIE, TranSalNet-Dense, TranSalNet-Res, UNISAL) and two classical image-statistics detectors (Spectral Residual, Fine-Grained).
  • Central-baseline dominance
    • A centered Gaussian (no training) raw AUC = 0.715 on the news photos; DeepGazeIIE = 0.692; other deep nets 0.682–0.692; classical detectors much lower (~0.55–0.57).
    • The centered Gaussian therefore outperformed every trained model on these real-world stimuli.
    • The centered Gaussian already captures ~92% of the predictive power of a held-out panel of real viewers; the best network covers ~82%.
  • What models actually learn
    • Regression of saliency maps onto simple predictors shows deep models heavily weight proximity to the image center (coefficients ~0.42–0.55). Classical detectors weight local contrast instead.
    • Inter-model correlations on the same photographs are low (r ≈ 0.14–0.31), indicating they capture similar but weak photo-specific signals.
  • Little beyond center
    • Once the central bias is removed (center-corrected/shuffled AUC), models add only ~3–5 percentage points of predictive accuracy (scores ≈ 0.51–0.52). On laboratory benchmarks they add far more.
    • A panel of real viewers grows to similar performance quickly: a panel of a few dozen viewers matches top models; the paper reports ~13 viewers suffice in some contexts and shows 256 viewers reach a performance plateau around the best model.
  • Heterogeneous fit and bias
    • The remaining model signal transfers unevenly across demographic/ideological groups. Models fit younger, White, and moderate viewers better; for Black viewers and ideologically extreme viewers the group-specific signal often falls below noise, leaving nothing for a model to learn.
    • The author proposes a learnability criterion: a group is learnable only if its members’ gaze patterns are internally consistent (and distinct from outsiders) on the images in question.
  • Robustness and controls
    • Pipelines reproduce standard laboratory benchmarks (e.g., MIT1003) where the networks perform strongly (DeepGazeIIE raw AUC ≈ 0.905), ruling out simple pipeline bugs.
    • The webcam measurement error was quantified (median ~3.7° visual angle) and tested; sensor/coarseness does not explain the central-bias dominance or the transfer failures.

Data & Methods

  • Recruitment and apparatus
    • National-quotas sampling on multiple demographic axes; participants used their own desktops/tablets/phones with webcam-based gaze estimation (streaming samples and dispersion-based fixation extraction).
    • Fixed-effects controls for recruitment wave and device; SEs clustered by participant; specification battery reported.
  • Scoring metrics
    • Raw AUC: probability that the model’s map value at a recorded gaze location exceeds map value at a uniformly random frame location (what buyers would see).
    • Center-corrected (shuffled AUC): comparison points drawn from gaze on other images to remove universal central bias and measure photo-specific content learned by the model.
    • Panels: held-out audience baselines built by pooling real viewers of the same photo and evaluated on viewers excluded from the pool.
  • Model diagnostics
    • Regress saliency map values (location-by-location, pooled over images) on four predictors: distance-to-center, person-detection indicator, local luminance contrast, and object size. This isolates what each model emphasizes.
  • Robustness checks
    • Replication of lab benchmark signals using identical scoring pipeline on MIT1003.
    • Assessed webcam noise and tested whether degrading lab gaze reproduces results — it did not.
    • All group comparisons use center-corrected scoring to remove central-bias confounds.

Implications for AI Economics

  • Valuation and market claims
    • Firms selling single-map predicted-attention products should be audited against a central baseline. The study suggests much of commercial value attributed to saliency models may be replaceable by trivial central-bias heuristics for many circulating images.
    • Market pricing that assumes fine-grained, universal human-attention prediction is likely overstated unless vendors can demonstrate predictive gains beyond center for the target audience.
  • Procurement and due diligence
    • Buyers (advertisers, platform designers, interface teams, clinical tool vendors) should require center-corrected performance metrics and group-level performance breakdowns before deploying model outputs as substitutes for real viewers.
    • Require demonstration of learnability for target subpopulations (i.e., show within-group gaze consistency distinct from outsiders) as a precondition for using conditional/personalized saliency.
  • Equity and regulatory risk
    • Because the model margins favor some demographic/ideological groups, downstream products that rely on single saliency maps (cropping, ranking, ad placement, interface prioritization) risk systematic disparate impacts. Economic analyses of platform effects must account for such heterogeneity.
    • Regulators and auditors should demand demographic-disaggregated evaluations and require that claims of "human-like" or "universal" attention be substantiated by group-level evidence.
  • Cost-effective alternatives and business practice
    • Collecting small panels of real viewers (the paper shows modest n often suffices) — enabled via webcam infrastructure — can match or beat model performance at low marginal cost per image and across groups.
    • For many commercial decisions, the cheapest reliable option may be a small, demographically appropriate panel rather than opaque general-purpose saliency predictions.
  • Modeling and research consequences
    • Economists studying attention-driven platform incentives, targeting, or welfare effects should incorporate heterogeneity in attention (not collapse to a single map). Using saliency-network outputs without group checks can produce biased estimates of who sees what and the downstream responses (clicks, votes, purchases).
    • The learnability criterion offers a principled way to decide when to invest in group-conditional models versus collecting small panels.
  • Recommendations (operational)
    • Mandate reporting of both raw and center-corrected AUC for any saliency product.
    • Benchmark models against: (i) centered Gaussian baseline; (ii) held-out audience panel baseline; (iii) demographic subgroup performance and learnability tests.
    • Prefer panel-based or personalized solutions where group-level signals are weak or where fairness across observed groups is required.

If you want, I can extract a compact checklist buyers or regulators can use when auditing saliency products (metrics to request, subgroup tests, sample sizes for verification panels).

Assessment

Paper Typedescriptive Evidence Strengthhigh — Large-scale, population-quota sample (3,023 US adults), 11.4 million raw webcam gaze samples across 61,458 viewings, multiple models tested, center-corrected and raw scoring conventions, held-out audience benchmarks, and extensive robustness checks (device and recording-error tests, replication on an independent public dataset) all support the empirical claims. Methods Rigorhigh — Careful sampling to national quotas, clustering SEs by participant, fixed effects for wave and device, two complementary gaze signals (raw samples and extracted fixations), multiple scoring conventions (raw and center-corrected AUC), held-out audience baselines, specification battery and replication on another eye-tracking dataset; the paper explicitly addresses plausible measurement confounds (webcam error, device, viewing time). Sample3,023 US adults recruited to national quotas on age, gender, education, race, income, and partisanship; they viewed 83 circulating Getty Images news photographs on their own desktops, tablets, and smartphones; dataset comprises 11.4 million webcam gaze samples from 61,458 viewings (with fixations extracted), and some analyses replicated on an independent public eye-tracking dataset. Themeshuman_ai_collab inequality adoption governance productivity GeneralizabilityStimuli restricted to 83 circulating news photographs (Getty Images) — may not generalize to other image genres (product shots, ads, UI screenshots, natural scenes, video)., Sample limited to US adults — cultural and geographic differences in gaze patterns not tested., Webcam-based gaze has coarser spatial precision than lab eye trackers (median error reported), so effects may differ in tightly controlled lab viewing., Tested a subset of available saliency models (four trained deep nets and two classical detectors); future models or fine-tuned systems might behave differently., Findings about single-map deployments may not apply to systems that personalize in real time or use multimodal/contextual signals beyond static saliency maps.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A blank centered Gaussian saliency map outperforms every trained deep network and both classical saliency detectors on the 83 circulating news photographs. Decision Quality positive Agreement between predicted saliency and recorded viewer gaze, measured by raw AUC
Reading fidelity high
Study strength high
n=3023
Central Gaussian AUC 0.715 versus 0.692 for the best trained model
0.3
The central Gaussian captures approximately 92% of the predictive performance achieved by a held-out audience of real viewers, while the strongest trained network captures approximately 82%. Decision Quality positive Predictive agreement with held-out gaze on news photographs
Reading fidelity high
Study strength high
n=3023
Central Gaussian 92% of held-out audience performance; strongest network 82%
0.3
The deep saliency networks add only a small amount of image-specific predictive information beyond the shared tendency of viewers to look near the center of the frame. Decision Quality positive Center-corrected agreement between saliency maps and gaze
Reading fidelity high
Study strength high
n=3023
Deep-network margins over center map: +0.035 to +0.044 AUC
0.3
The content learned by DeepGazeIIE on top of the central tendency is misaligned with the gaze of the audiences viewing these news photographs. Decision Quality negative Pairwise ranking accuracy of predicted saliency relative to actual gaze
Reading fidelity high
Study strength high
n=3023
DeepGazeIIE lost 0.037 of raw score on crossing pairs and recovered 0.021 on mirror pairs
0.3
The saliency networks perform substantially better on the MIT1003 laboratory benchmark than on the circulating news photographs viewed through webcams. Decision Quality positive Raw AUC agreement between saliency maps and gaze
Reading fidelity high
Study strength medium
DeepGazeIIE AUC 0.905 on MIT1003 versus 0.692 on news photographs
0.18
A panel of real viewers who viewed the same photograph predicts a new viewer's gaze nearly as well as the best saliency model, and increasing the panel to 256 viewers provides little additional predictive performance. Decision Quality positive Center-corrected predictive agreement with held-out viewers' gaze
Reading fidelity high
Study strength high
n=256
Panel score 0.550 at 256 viewers versus 0.553 for UNISAL
0.3
The predictive accuracy of the saliency models is systematically higher for younger, White, and politically moderate viewers than for older, Black, and ideologically extreme viewers. Inequality mixed Group-specific center-corrected agreement between saliency predictions and viewer gaze
Reading fidelity high
Study strength medium
n=3023
0.18
Younger audiences exhibit a reproducible group-specific visual-attention signature that can be learned and generalized to photographs not used for training. Decision Quality positive Generalization of group-specific gaze patterns across images and datasets
Reading fidelity high
Study strength medium
Signature reappeared at twice the size in an independent dataset
0.18
For Black viewers and ideologically extreme viewers, the study finds no detectable group-specific gaze signal beyond their standard errors in either of the two gaze measures. Decision Quality null_result Detectable, reproducible group-specific visual-attention signal
Reading fidelity high
Study strength medium
Signals fell below their own standard errors on both gaze measures
0.18
The study's webcam-based gaze measurement is substantially noisier than laboratory measurement, but the difference in model performance cannot be reproduced by manipulating measurement error alone. Decision Quality null_result Effect of gaze-measurement error on saliency-model performance
Reading fidelity high
Study strength medium
n=3023
Median recording error 3.68° for webcam gaze versus 0.76° for laboratory gaze
0.18

Notes