The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI assistants largely ignore local eateries: a census audit in two Bali submarkets finds at least 85.6% of cafes, restaurants and bars never surface in four production systems' recommendations, with discoverability driven by documentation (reviews, websites, price listings) while star ratings only influence rank among those already recommended; closed venues are recommended far more often than invented ones.

Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census
Vladimir Pitenin · August 07, 2026
arxiv descriptive high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Vladimir Pitenin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Vladimir Pitenin provider ID
In a pre-registered census audit of 4,776 Bali food-and-drink venues, four production AI assistants failed to recommend the vast majority (≥85.6%) of venues; entry into answers correlates with documentation signals (review volume, website, price, web mentions) whereas star rating predicts relative rank among already-recommended venues, and staleness (closed venues) is a more common failure than fabrication.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.

Summary

Main Finding

AI assistants systematically omit the large majority of local food-and-drink venues: in a complete census of 4,776 cafés, restaurants, and bars in two Bali submarkets, 85.6% of venues were never recommended by any of four production AI systems (ChatGPT, Claude, Gemini, Perplexity) across 2,208 query runs — with even established venues (≥50 ratings) experiencing 72.6% invisibility. Visibility operates on two margins: documentation (reviews, website, price, web mentions) predicts whether a venue is ever recommended, while star rating predicts rank among venues that are recommended.

Key Points

  • Scope and scale
    • Complete market census: 4,776 food-and-drink venues in greater Canggu and greater Ubud, Bali.
    • Audit: 2,208 runs (96 persona-conditioned queries × systems × repeats), producing 12,439 valid venue mentions and ~1.4 million venue-level exposure opportunities.
    • Systems audited: OpenAI (gpt-5.2 via search-grounded Responses API), Anthropic (claude-sonnet-5 with web tool), Google (gemini-3.5-flash grounded to Search), Perplexity (sonar).
  • Invisiblity and concentration
    • 85.6% of venues never recommended by any system during the wave; under broader defensible market framing the floor exceeds 92%.
    • Even among venues with ≥50 reviews, 72.6% were never recommended.
    • Visibility is long-tailed; the single most-recommended venue accounts for only 1.9% of recommendations.
  • Two-margin structure (entry vs. rank)
    • Entry (whether a venue appears at all) is associated with documentation signals:
      • Review volume: OR 1.64
      • Has an own website: OR 1.92
      • Listed price information: OR 1.54
      • Third‑party web mentions: OR 1.44
      • Star rating: null at entry (OR 0.89)
    • Rank within recommended venues:
      • Higher star rating predicts occupying top positions (OR 1.17 for first position).
  • Other empirical findings
    • Presence in an open POI dataset (Foursquare) shows no positive effect on either entry or rank once documentation is controlled.
    • Fabrication is rare (≈0.08% of mentions likely invented), but staleness is a tangible practical failure: 93 recommendations were for permanently closed venues.
    • Cross-system agreement is low (top‑20 Jaccard index between systems ≈ 0.33–0.54).
    • A two-week test–retest indicates run-to-run churn is comparable to same-day reruns; visibility appears to be a persistent property measured with stochastic sampling noise.
  • Methodological rigor
    • Pre-registered protocol; adaptive Places API grid to avoid truncated enumeration; double-annotation for extraction and entity matching; matching-error remediation changed one factor’s effect from significantly positive to null — showing sensitivity to matching quality.
    • Census is Google Places–based and augmented by resolving unmatched AI-recommended names; reported invisibility rates are conservative floors.

Data & Methods

  • Population frame
    • Defined by Google Places listings of five food-service types inside fixed polygons for greater Canggu and greater Ubud.
    • Adaptive grid over Places Nearby Search API to avoid 20-result truncation; recursive subdivision down to 130 m cells.
    • Augmented with targeted Place Text Search probes for AI-recommended names that failed to match the grid, producing a final registry of 4,776 venues (each with Google snapshot metadata: rating, review count, price level, hours, website, business status).
    • Capture–recapture and hand-audit against Foursquare used to assess coverage; estimates make invisibility measures conservative.
  • Query instrument
    • 8 personas × 6 paraphrase templates × 2 areas = 96 unique persona-conditioned, first-person discovery prompts (e.g., remote-worker café, date night, budget traveler).
    • Prompts designed, frozen, and pre-registered prior to confirmatory collection; paraphrase sensitivity examined.
  • Systems & collection
    • Four production, search-grounded assistants queried via public/grounding APIs; runs spread across seven days.
    • Repeats per query: Perplexity (10), OpenAI & Gemini (5), Claude (3) — unequal repeats handled in estimation; a two-week test–retest subset (holdout) also collected.
    • Extraction and entity matching validated with independent double annotation; matching errors measured and corrected in confirmatory analysis.
  • Outcome coding
    • Mentions classified as recommended, neutral-mentioned, or advised-against; unmatched mentions examined to detect fabrication vs. stale/closed venues.
  • Statistical analysis
    • Logistic models for entry (any mention) and conditional models for rank within answers.
    • Odds ratios reported for key predictors; sensitivity checks including removing Claude arm (due to search-cap limit) and re-running after matching remediation.

Implications for AI Economics

  • Distributional revenue consequences
    • Given prior causal estimates that discovery visibility and displayed ratings materially affect restaurant revenue, the high AI invisibility rate implies substantial and asymmetric demand allocation driven by AI intermediaries. Many small, independent operators may lose discoverable demand entirely in AI-mediated channels.
  • Visibility is infrastructure-driven, not purely reputation-driven
    • Documentation signals (website, review volume, price listings, web mentions) matter for entry. This suggests investments in basic online infrastructure and third‑party mentions are likely to increase the chance of being included in AI recommendations — a practical GEO (generative engine optimization) lever for small businesses.
    • Conversely, presence in open POI datasets (e.g., Foursquare) alone is not sufficient; platform-specific and web-documentation signals appear more consequential.
  • Policy and market design
    • Low cross-system agreement and high invisibility raise concerns about gatekeeping by opaque generative pipelines. Regulators and platform designers should consider transparency, auditability, and updating cadence (to avoid staleness) as design priorities.
    • Remedies could include standardized, machine-readable POI feeds (with frequent updates), minimum disclosure of grounding sources, or tools enabling venues to signal up-to-date documentation to generative systems.
  • Methodological guidance for future audits and industry tracking
    • Census-denominated audits reveal margins that sample-based or catalogue audits miss (the entry margin vs. rank margin). Audits of discovery systems should (a) validate entity matching carefully, (b) use bounded market censuses where feasible, and (c) distinguish documentation vs. reputation effects.
    • The practical failure mode (staleness) implies that improving POI freshness and grounding quality is a high-impact lever for both platforms and venues.
  • Commercial strategy implications
    • For venue owners and service providers (GEO vendors): prioritize review acquisition, maintained websites, explicit price metadata, and third-party coverage; investments in open-POI listing alone may not move the needle.
    • For intermediaries and platforms: marketplace power and attention allocation in AI outputs represent an economic externality; designing smaller businesses’ access to the discovery pipeline could materially alter local market competition.

(Replication materials, the census-construction method, the pre-registered protocol, and derived data are released with the paper.)

Assessment

Paper Typedescriptive Evidence Strengthhigh — A pre-registered, census-based audit with full enumeration of 4,776 venues, multi-system sampling (2,208 runs), validated extraction and double annotation, and sensitivity/test-retest checks provides strong descriptive evidence about what these production AI systems surface within the studied markets; causal claims are not made. Methods Rigorhigh — Carefully constructed population frame (adaptive Google Places grid + resolution passes), pre-registration, repeated runs and paraphrases, multi-system coverage, independent validation of entity matching, and pre-registered models and sensitivity analyses support rigorous measurement; limitations include observational design, two-market scope, and some vendor-specific constraints (e.g., Claude search cap). SampleA complete census of 4,776 food-and-drink venues (cafes, restaurants, bars, coffee shops, bakeries) within two bounded Bali submarkets (greater Canggu and greater Ubud), built from Google Places via an adaptive grid and augmented by AI-resolved probes; audit data comprise 2,208 confirmatory runs across four production search-grounded assistants (OpenAI gpt-5.2, Anthropic Claude, Google Gemini, Perplexity), 96 persona-conditioned queries repeated over seven days (plus pilot and two-week retest), yielding 12,439 valid venue mentions and ~1.4 million venue-level exposure opportunities. Themesadoption inequality GeneralizabilityGeographic limitation: two Bali submarkets (tourist-oriented, high density of independent venues) may not reflect larger or non-tourist metros., Language/Persona scope: English-language, traveler/nomad-style persona templates may bias which venues are surfaced compared with local-language or other user segments., Temporal snapshot: seven-day primary wave (and two-week retest) captures system behaviour at specific timestamps and model/config states that may change over time., API/interface limitation: analysis restricted to search-grounded public APIs and vendor consumer-default configurations (Claude search cap, Gemini grounding constraints), not internal/platform-level ranking signals., Frame dependence: census relies on Google Places frame; venues absent from that frame are treated as never recommended but may exist in other listings.

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
At least 85.6% of venues in the two-market census were never recommended by any audited AI system in any run. Adoption Rate negative Whether a venue was surfaced in AI recommendations
Reading fidelity high
Study strength high
n=4776
85.6% never recommended
0.3
Among established venues with at least 50 ratings, 72.6% were never recommended by any audited AI system. Adoption Rate negative Whether an established venue was surfaced in AI recommendations
Reading fidelity high
Study strength high
n=4776
72.6% never recommended
0.3
Venue entry into AI answers was positively associated with review volume, with an odds ratio of 1.64. Adoption Rate positive Venue entry into an AI-generated recommendation answer
Reading fidelity high
Study strength medium
n=4776
OR 1.64
0.18
Having an own website was positively associated with venue entry into AI answers, with an odds ratio of 1.92. Adoption Rate positive Venue entry into an AI-generated recommendation answer
Reading fidelity high
Study strength medium
n=4776
OR 1.92
0.18
Listed price information was positively associated with venue entry into AI answers, with an odds ratio of 1.54. Adoption Rate positive Venue entry into an AI-generated recommendation answer
Reading fidelity high
Study strength medium
n=4776
OR 1.54
0.18
Third-party web mentions were positively associated with venue entry into AI answers, with an odds ratio of 1.44. Adoption Rate positive Venue entry into an AI-generated recommendation answer
Reading fidelity high
Study strength medium
n=4776
OR 1.44
0.18
Star rating had no positive association with whether a venue entered an AI answer once review volume was controlled, with an odds ratio of 0.89. Adoption Rate null_result Venue entry into an AI-generated recommendation answer
Reading fidelity high
Study strength medium
n=4776
OR 0.89
0.18
Among venues that were recommended, rating significantly predicted appearing in the first position, with an odds ratio of 1.17. Adoption Rate positive Whether a recommended venue appeared in first position
Reading fidelity high
Study strength medium
OR 1.17
0.18
Presence in the Foursquare open point-of-interest dataset showed no positive effect on AI recommendation visibility at either the entry or ranking margin after documentation controls. Adoption Rate null_result AI recommendation entry and within-answer ranking
Reading fidelity high
Study strength medium
n=4776
0.18
Outright fabrication was rare: one likely-invented venue name occurred among more than 12,000 valid venue mentions, corresponding to 0.08% of mentions. Error Rate negative Rate of fabricated or unmatched venue mentions
Reading fidelity high
Study strength medium
n=12439
0.08% of mentions
0.18
The audited systems recommended permanently closed venues 93 times, making staleness more common than outright fabrication. Error Rate negative Recommendations of permanently closed venues
Reading fidelity high
Study strength medium
n=12439
93 recommendations
0.18
Cross-system agreement on the top 20 recommended venues was low, with Jaccard similarity ranging from 0.33 to 0.54. Market Structure mixed Overlap in top-20 venue recommendation sets across AI systems
Reading fidelity high
Study strength medium
n=4
top-20 Jaccard 0.33–0.54
0.18
A two-week test–retest produced cross-period answer similarity comparable to same-day rerun similarity, indicating no temporal drift beyond the systems' sampling noise. Adoption Rate null_result Similarity of AI recommendation answers across time
Reading fidelity high
Study strength medium
n=144
0.18

Notes