0 cumulative citations
View corpus contextAI assistants largely ignore local eateries: a census audit in two Bali submarkets finds at least 85.6% of cafes, restaurants and bars never surface in four production systems' recommendations, with discoverability driven by documentation (reviews, websites, price listings) while star ratings only influence rank among those already recommended; closed venues are recommended far more often than invented ones.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.
Summary
Main Finding
AI assistants systematically omit the large majority of local food-and-drink venues: in a complete census of 4,776 cafés, restaurants, and bars in two Bali submarkets, 85.6% of venues were never recommended by any of four production AI systems (ChatGPT, Claude, Gemini, Perplexity) across 2,208 query runs — with even established venues (≥50 ratings) experiencing 72.6% invisibility. Visibility operates on two margins: documentation (reviews, website, price, web mentions) predicts whether a venue is ever recommended, while star rating predicts rank among venues that are recommended.
Key Points
- Scope and scale
- Complete market census: 4,776 food-and-drink venues in greater Canggu and greater Ubud, Bali.
- Audit: 2,208 runs (96 persona-conditioned queries × systems × repeats), producing 12,439 valid venue mentions and ~1.4 million venue-level exposure opportunities.
- Systems audited: OpenAI (gpt-5.2 via search-grounded Responses API), Anthropic (claude-sonnet-5 with web tool), Google (gemini-3.5-flash grounded to Search), Perplexity (sonar).
- Invisiblity and concentration
- 85.6% of venues never recommended by any system during the wave; under broader defensible market framing the floor exceeds 92%.
- Even among venues with ≥50 reviews, 72.6% were never recommended.
- Visibility is long-tailed; the single most-recommended venue accounts for only 1.9% of recommendations.
- Two-margin structure (entry vs. rank)
- Entry (whether a venue appears at all) is associated with documentation signals:
- Review volume: OR 1.64
- Has an own website: OR 1.92
- Listed price information: OR 1.54
- Third‑party web mentions: OR 1.44
- Star rating: null at entry (OR 0.89)
- Rank within recommended venues:
- Higher star rating predicts occupying top positions (OR 1.17 for first position).
- Entry (whether a venue appears at all) is associated with documentation signals:
- Other empirical findings
- Presence in an open POI dataset (Foursquare) shows no positive effect on either entry or rank once documentation is controlled.
- Fabrication is rare (≈0.08% of mentions likely invented), but staleness is a tangible practical failure: 93 recommendations were for permanently closed venues.
- Cross-system agreement is low (top‑20 Jaccard index between systems ≈ 0.33–0.54).
- A two-week test–retest indicates run-to-run churn is comparable to same-day reruns; visibility appears to be a persistent property measured with stochastic sampling noise.
- Methodological rigor
- Pre-registered protocol; adaptive Places API grid to avoid truncated enumeration; double-annotation for extraction and entity matching; matching-error remediation changed one factor’s effect from significantly positive to null — showing sensitivity to matching quality.
- Census is Google Places–based and augmented by resolving unmatched AI-recommended names; reported invisibility rates are conservative floors.
Data & Methods
- Population frame
- Defined by Google Places listings of five food-service types inside fixed polygons for greater Canggu and greater Ubud.
- Adaptive grid over Places Nearby Search API to avoid 20-result truncation; recursive subdivision down to 130 m cells.
- Augmented with targeted Place Text Search probes for AI-recommended names that failed to match the grid, producing a final registry of 4,776 venues (each with Google snapshot metadata: rating, review count, price level, hours, website, business status).
- Capture–recapture and hand-audit against Foursquare used to assess coverage; estimates make invisibility measures conservative.
- Query instrument
- 8 personas × 6 paraphrase templates × 2 areas = 96 unique persona-conditioned, first-person discovery prompts (e.g., remote-worker café, date night, budget traveler).
- Prompts designed, frozen, and pre-registered prior to confirmatory collection; paraphrase sensitivity examined.
- Systems & collection
- Four production, search-grounded assistants queried via public/grounding APIs; runs spread across seven days.
- Repeats per query: Perplexity (10), OpenAI & Gemini (5), Claude (3) — unequal repeats handled in estimation; a two-week test–retest subset (holdout) also collected.
- Extraction and entity matching validated with independent double annotation; matching errors measured and corrected in confirmatory analysis.
- Outcome coding
- Mentions classified as recommended, neutral-mentioned, or advised-against; unmatched mentions examined to detect fabrication vs. stale/closed venues.
- Statistical analysis
- Logistic models for entry (any mention) and conditional models for rank within answers.
- Odds ratios reported for key predictors; sensitivity checks including removing Claude arm (due to search-cap limit) and re-running after matching remediation.
Implications for AI Economics
- Distributional revenue consequences
- Given prior causal estimates that discovery visibility and displayed ratings materially affect restaurant revenue, the high AI invisibility rate implies substantial and asymmetric demand allocation driven by AI intermediaries. Many small, independent operators may lose discoverable demand entirely in AI-mediated channels.
- Visibility is infrastructure-driven, not purely reputation-driven
- Documentation signals (website, review volume, price listings, web mentions) matter for entry. This suggests investments in basic online infrastructure and third‑party mentions are likely to increase the chance of being included in AI recommendations — a practical GEO (generative engine optimization) lever for small businesses.
- Conversely, presence in open POI datasets (e.g., Foursquare) alone is not sufficient; platform-specific and web-documentation signals appear more consequential.
- Policy and market design
- Low cross-system agreement and high invisibility raise concerns about gatekeeping by opaque generative pipelines. Regulators and platform designers should consider transparency, auditability, and updating cadence (to avoid staleness) as design priorities.
- Remedies could include standardized, machine-readable POI feeds (with frequent updates), minimum disclosure of grounding sources, or tools enabling venues to signal up-to-date documentation to generative systems.
- Methodological guidance for future audits and industry tracking
- Census-denominated audits reveal margins that sample-based or catalogue audits miss (the entry margin vs. rank margin). Audits of discovery systems should (a) validate entity matching carefully, (b) use bounded market censuses where feasible, and (c) distinguish documentation vs. reputation effects.
- The practical failure mode (staleness) implies that improving POI freshness and grounding quality is a high-impact lever for both platforms and venues.
- Commercial strategy implications
- For venue owners and service providers (GEO vendors): prioritize review acquisition, maintained websites, explicit price metadata, and third-party coverage; investments in open-POI listing alone may not move the needle.
- For intermediaries and platforms: marketplace power and attention allocation in AI outputs represent an economic externality; designing smaller businesses’ access to the discovery pipeline could materially alter local market competition.
(Replication materials, the census-construction method, the pre-registered protocol, and derived data are released with the paper.)
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| At least 85.6% of venues in the two-market census were never recommended by any audited AI system in any run. Adoption Rate | negative | Whether a venue was surfaced in AI recommendations |
Reading fidelity
high
Study strength
high
|
n=4776
85.6% never recommended
|
| Among established venues with at least 50 ratings, 72.6% were never recommended by any audited AI system. Adoption Rate | negative | Whether an established venue was surfaced in AI recommendations |
Reading fidelity
high
Study strength
high
|
n=4776
72.6% never recommended
|
| Venue entry into AI answers was positively associated with review volume, with an odds ratio of 1.64. Adoption Rate | positive | Venue entry into an AI-generated recommendation answer |
Reading fidelity
high
Study strength
medium
|
n=4776
OR 1.64
|
| Having an own website was positively associated with venue entry into AI answers, with an odds ratio of 1.92. Adoption Rate | positive | Venue entry into an AI-generated recommendation answer |
Reading fidelity
high
Study strength
medium
|
n=4776
OR 1.92
|
| Listed price information was positively associated with venue entry into AI answers, with an odds ratio of 1.54. Adoption Rate | positive | Venue entry into an AI-generated recommendation answer |
Reading fidelity
high
Study strength
medium
|
n=4776
OR 1.54
|
| Third-party web mentions were positively associated with venue entry into AI answers, with an odds ratio of 1.44. Adoption Rate | positive | Venue entry into an AI-generated recommendation answer |
Reading fidelity
high
Study strength
medium
|
n=4776
OR 1.44
|
| Star rating had no positive association with whether a venue entered an AI answer once review volume was controlled, with an odds ratio of 0.89. Adoption Rate | null_result | Venue entry into an AI-generated recommendation answer |
Reading fidelity
high
Study strength
medium
|
n=4776
OR 0.89
|
| Among venues that were recommended, rating significantly predicted appearing in the first position, with an odds ratio of 1.17. Adoption Rate | positive | Whether a recommended venue appeared in first position |
Reading fidelity
high
Study strength
medium
|
OR 1.17
|
| Presence in the Foursquare open point-of-interest dataset showed no positive effect on AI recommendation visibility at either the entry or ranking margin after documentation controls. Adoption Rate | null_result | AI recommendation entry and within-answer ranking |
Reading fidelity
high
Study strength
medium
|
n=4776
|
| Outright fabrication was rare: one likely-invented venue name occurred among more than 12,000 valid venue mentions, corresponding to 0.08% of mentions. Error Rate | negative | Rate of fabricated or unmatched venue mentions |
Reading fidelity
high
Study strength
medium
|
n=12439
0.08% of mentions
|
| The audited systems recommended permanently closed venues 93 times, making staleness more common than outright fabrication. Error Rate | negative | Recommendations of permanently closed venues |
Reading fidelity
high
Study strength
medium
|
n=12439
93 recommendations
|
| Cross-system agreement on the top 20 recommended venues was low, with Jaccard similarity ranging from 0.33 to 0.54. Market Structure | mixed | Overlap in top-20 venue recommendation sets across AI systems |
Reading fidelity
high
Study strength
medium
|
n=4
top-20 Jaccard 0.33–0.54
|
| A two-week test–retest produced cross-period answer similarity comparable to same-day rerun similarity, indicating no temporal drift beyond the systems' sampling noise. Adoption Rate | null_result | Similarity of AI recommendation answers across time |
Reading fidelity
high
Study strength
medium
|
n=144
|