0 cumulative citations
View corpus contextPredictive saliency models miss much of what audiences actually look at: across a nationally quotaed sample viewing news photos, a simple centered map outperforms leading trained models, and model errors disproportionately affect older, Black, and politically extreme viewers, raising fairness risks for systems that use a single attention map.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see everyone, and this study supplies the standard by which such a claim should be judged.
Summary
Main Finding
On circulating news photographs viewed in situ by a demographically representative US sample (webcam gaze), commercial and research saliency networks largely rely on a universal central bias. A blank centered Gaussian map (no training) predicts viewers’ gaze better than every trained network tested. The trained models add only a thin margin beyond that center, and that margin transfers unevenly across demographic groups — favoring younger, White, and politically moderate viewers and failing for older, Black, and ideologically extreme viewers. A small panel of real viewers of the same photo already matches or exceeds model performance for many use cases, suggesting practical remedies.
Key Points
- Data scale and sample
- 3,023 US adults recruited to national quotas (age, gender, education, race, income, partisanship).
- 11.4 million webcam gaze samples across 61,458 viewings of 83 circulating Getty news photographs.
- Models evaluated
- Six predictors: four deep saliency networks (DeepGazeIIE, TranSalNet-Dense, TranSalNet-Res, UNISAL) and two classical image-statistics detectors (Spectral Residual, Fine-Grained).
- Central-baseline dominance
- A centered Gaussian (no training) raw AUC = 0.715 on the news photos; DeepGazeIIE = 0.692; other deep nets 0.682–0.692; classical detectors much lower (~0.55–0.57).
- The centered Gaussian therefore outperformed every trained model on these real-world stimuli.
- The centered Gaussian already captures ~92% of the predictive power of a held-out panel of real viewers; the best network covers ~82%.
- What models actually learn
- Regression of saliency maps onto simple predictors shows deep models heavily weight proximity to the image center (coefficients ~0.42–0.55). Classical detectors weight local contrast instead.
- Inter-model correlations on the same photographs are low (r ≈ 0.14–0.31), indicating they capture similar but weak photo-specific signals.
- Little beyond center
- Once the central bias is removed (center-corrected/shuffled AUC), models add only ~3–5 percentage points of predictive accuracy (scores ≈ 0.51–0.52). On laboratory benchmarks they add far more.
- A panel of real viewers grows to similar performance quickly: a panel of a few dozen viewers matches top models; the paper reports ~13 viewers suffice in some contexts and shows 256 viewers reach a performance plateau around the best model.
- Heterogeneous fit and bias
- The remaining model signal transfers unevenly across demographic/ideological groups. Models fit younger, White, and moderate viewers better; for Black viewers and ideologically extreme viewers the group-specific signal often falls below noise, leaving nothing for a model to learn.
- The author proposes a learnability criterion: a group is learnable only if its members’ gaze patterns are internally consistent (and distinct from outsiders) on the images in question.
- Robustness and controls
- Pipelines reproduce standard laboratory benchmarks (e.g., MIT1003) where the networks perform strongly (DeepGazeIIE raw AUC ≈ 0.905), ruling out simple pipeline bugs.
- The webcam measurement error was quantified (median ~3.7° visual angle) and tested; sensor/coarseness does not explain the central-bias dominance or the transfer failures.
Data & Methods
- Recruitment and apparatus
- National-quotas sampling on multiple demographic axes; participants used their own desktops/tablets/phones with webcam-based gaze estimation (streaming samples and dispersion-based fixation extraction).
- Fixed-effects controls for recruitment wave and device; SEs clustered by participant; specification battery reported.
- Scoring metrics
- Raw AUC: probability that the model’s map value at a recorded gaze location exceeds map value at a uniformly random frame location (what buyers would see).
- Center-corrected (shuffled AUC): comparison points drawn from gaze on other images to remove universal central bias and measure photo-specific content learned by the model.
- Panels: held-out audience baselines built by pooling real viewers of the same photo and evaluated on viewers excluded from the pool.
- Model diagnostics
- Regress saliency map values (location-by-location, pooled over images) on four predictors: distance-to-center, person-detection indicator, local luminance contrast, and object size. This isolates what each model emphasizes.
- Robustness checks
- Replication of lab benchmark signals using identical scoring pipeline on MIT1003.
- Assessed webcam noise and tested whether degrading lab gaze reproduces results — it did not.
- All group comparisons use center-corrected scoring to remove central-bias confounds.
Implications for AI Economics
- Valuation and market claims
- Firms selling single-map predicted-attention products should be audited against a central baseline. The study suggests much of commercial value attributed to saliency models may be replaceable by trivial central-bias heuristics for many circulating images.
- Market pricing that assumes fine-grained, universal human-attention prediction is likely overstated unless vendors can demonstrate predictive gains beyond center for the target audience.
- Procurement and due diligence
- Buyers (advertisers, platform designers, interface teams, clinical tool vendors) should require center-corrected performance metrics and group-level performance breakdowns before deploying model outputs as substitutes for real viewers.
- Require demonstration of learnability for target subpopulations (i.e., show within-group gaze consistency distinct from outsiders) as a precondition for using conditional/personalized saliency.
- Equity and regulatory risk
- Because the model margins favor some demographic/ideological groups, downstream products that rely on single saliency maps (cropping, ranking, ad placement, interface prioritization) risk systematic disparate impacts. Economic analyses of platform effects must account for such heterogeneity.
- Regulators and auditors should demand demographic-disaggregated evaluations and require that claims of "human-like" or "universal" attention be substantiated by group-level evidence.
- Cost-effective alternatives and business practice
- Collecting small panels of real viewers (the paper shows modest n often suffices) — enabled via webcam infrastructure — can match or beat model performance at low marginal cost per image and across groups.
- For many commercial decisions, the cheapest reliable option may be a small, demographically appropriate panel rather than opaque general-purpose saliency predictions.
- Modeling and research consequences
- Economists studying attention-driven platform incentives, targeting, or welfare effects should incorporate heterogeneity in attention (not collapse to a single map). Using saliency-network outputs without group checks can produce biased estimates of who sees what and the downstream responses (clicks, votes, purchases).
- The learnability criterion offers a principled way to decide when to invest in group-conditional models versus collecting small panels.
- Recommendations (operational)
- Mandate reporting of both raw and center-corrected AUC for any saliency product.
- Benchmark models against: (i) centered Gaussian baseline; (ii) held-out audience panel baseline; (iii) demographic subgroup performance and learnability tests.
- Prefer panel-based or personalized solutions where group-level signals are weak or where fairness across observed groups is required.
If you want, I can extract a compact checklist buyers or regulators can use when auditing saliency products (metrics to request, subgroup tests, sample sizes for verification panels).
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| A blank centered Gaussian saliency map outperforms every trained deep network and both classical saliency detectors on the 83 circulating news photographs. Decision Quality | positive | Agreement between predicted saliency and recorded viewer gaze, measured by raw AUC |
Reading fidelity
high
Study strength
high
|
n=3023
Central Gaussian AUC 0.715 versus 0.692 for the best trained model
|
| The central Gaussian captures approximately 92% of the predictive performance achieved by a held-out audience of real viewers, while the strongest trained network captures approximately 82%. Decision Quality | positive | Predictive agreement with held-out gaze on news photographs |
Reading fidelity
high
Study strength
high
|
n=3023
Central Gaussian 92% of held-out audience performance; strongest network 82%
|
| The deep saliency networks add only a small amount of image-specific predictive information beyond the shared tendency of viewers to look near the center of the frame. Decision Quality | positive | Center-corrected agreement between saliency maps and gaze |
Reading fidelity
high
Study strength
high
|
n=3023
Deep-network margins over center map: +0.035 to +0.044 AUC
|
| The content learned by DeepGazeIIE on top of the central tendency is misaligned with the gaze of the audiences viewing these news photographs. Decision Quality | negative | Pairwise ranking accuracy of predicted saliency relative to actual gaze |
Reading fidelity
high
Study strength
high
|
n=3023
DeepGazeIIE lost 0.037 of raw score on crossing pairs and recovered 0.021 on mirror pairs
|
| The saliency networks perform substantially better on the MIT1003 laboratory benchmark than on the circulating news photographs viewed through webcams. Decision Quality | positive | Raw AUC agreement between saliency maps and gaze |
Reading fidelity
high
Study strength
medium
|
DeepGazeIIE AUC 0.905 on MIT1003 versus 0.692 on news photographs
|
| A panel of real viewers who viewed the same photograph predicts a new viewer's gaze nearly as well as the best saliency model, and increasing the panel to 256 viewers provides little additional predictive performance. Decision Quality | positive | Center-corrected predictive agreement with held-out viewers' gaze |
Reading fidelity
high
Study strength
high
|
n=256
Panel score 0.550 at 256 viewers versus 0.553 for UNISAL
|
| The predictive accuracy of the saliency models is systematically higher for younger, White, and politically moderate viewers than for older, Black, and ideologically extreme viewers. Inequality | mixed | Group-specific center-corrected agreement between saliency predictions and viewer gaze |
Reading fidelity
high
Study strength
medium
|
n=3023
|
| Younger audiences exhibit a reproducible group-specific visual-attention signature that can be learned and generalized to photographs not used for training. Decision Quality | positive | Generalization of group-specific gaze patterns across images and datasets |
Reading fidelity
high
Study strength
medium
|
Signature reappeared at twice the size in an independent dataset
|
| For Black viewers and ideologically extreme viewers, the study finds no detectable group-specific gaze signal beyond their standard errors in either of the two gaze measures. Decision Quality | null_result | Detectable, reproducible group-specific visual-attention signal |
Reading fidelity
high
Study strength
medium
|
Signals fell below their own standard errors on both gaze measures
|
| The study's webcam-based gaze measurement is substantially noisier than laboratory measurement, but the difference in model performance cannot be reproduced by manipulating measurement error alone. Decision Quality | null_result | Effect of gaze-measurement error on saliency-model performance |
Reading fidelity
high
Study strength
medium
|
n=3023
Median recording error 3.68° for webcam gaze versus 0.76° for laboratory gaze
|