The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Generative search favors local suppliers only when users ask in the local language: English queries on the same connection return global brands while exit IP determines which national market is referenced; the browser UI and API are equally unstable, so single snapshots can be misleading.

The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface
Dmitrij Żatuchin · August 30, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Dmitrij Żatuchin unresolved corpus identity

Semantic Scholar

Latest observation:

  1. D. Żatuchin provider ID
In a controlled probe, query language (not user IP/location) gates whether generative search names local suppliers—asking in the local language yields local recommendations while English returns global brands—and exit IP separately determines which national market is referenced, with substantial run-to-run instability across surfaces.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third, language and location are separable and act on different things: holding the query language fixed and moving only the exit IP moves the market whose brands are named while the answer stays in the query language. We show this on two unrelated pairs, Turkish asked from Berlin and Russian asked from Tallinn, and in both the answer names the resident country's suppliers. A minority language occupies a middle tier: Russian asked from Estonia names an Estonian supplier in 4 of 6 runs and a global one in all six, where Estonian names a local supplier in every run and English names none. A negative control in a second category, coded with the same instrument, shows no language effect at all, and disconfirms our own expectation: that category does have domestic suppliers and none was named in any language, which points the explanation at whether a category is nationally regulated rather than at whether it is nationally supplied.

Summary

Main Finding

A generative search interface’s commercial recommendations are strongly gated by the query language (which decides whether localisation is attempted) and separately by the exit IP (which decides which national market is treated as the user’s). Asked in a local language, the model overwhelmingly names local suppliers; asked in English from the same connection it names global suppliers. The top recommendation is also non-deterministic: identical runs often return different top picks.

Key Points

  • Language as gate: When the query language matched the country’s official language, local suppliers dominated (e.g., Estonian-language queries from Tallinn named Estonian accounting vendors in 6/6 runs). The same connection asked in English named global vendors (Estonia: local suppliers 6/6 in Estonian vs 0/6 in English).
  • Exit-IP market selection is separable from language: holding query language fixed and changing only the exit IP changed the market named while the answer stayed in the same language (example: Turkish queries from Istanbul → Turkish suppliers; Turkish queries from Berlin → German suppliers, still in Turkish).
  • Three-tier ordering on one connection: On the Tallinn connection, Estonian (official) > Russian (local minority) > English (global) in terms of producing local supplier mentions. Russian often produced a mix of local + global; English produced only global names.
  • Instability: Across six identical runs per prompt, the top recommendation changed on 4/6 prompts in both the logged-out web UI and in the API (with and without web search). Instability rates were equal across surfaces, implying system-level nondeterminism.
  • API citation behavior: API calls with web_search enabled returned citations (200 citations across 36 runs; ~5.6 per answer), while API calls with web_search disabled returned none. The API also lacked user geography; it returned U.S.-centric answers when asked in English unless exit IP was otherwise specified by the collector.
  • Negative control: For a second category (project-management software), domestic suppliers were not named in any language or exit-IP. This disconfirms the simple “exists-a-domestic-supplier → language effect” expectation and suggests category-specific factors (likely regulatory constraints) determine whether localisation is invoked.
  • Statistical support: Two independent replications (Germany and Norway) show a significant language effect for local-vs-English winners (combined Fisher’s method p ≈ 0.0022).

Data & Methods

  • Scope: 234 usable runs collected 29–30 Aug 2026 across 11 cells (combinations of surface, exit IP and query language); six identical runs per prompt per cell (plus some discarded/recollected runs for instrument checks).
  • Cells: exits included Berlin (DE), Oslo (NO), Istanbul (TR) and a residential Tallinn (EE) connection. Surface: logged-out ChatGPT web interface and OpenAI API (gpt-5.6-terra) with web_search on/off.
  • Prompts: Treatment = accounting software for freelancers (chosen because each country has domestic suppliers). Control = project management for a small marketing agency.
  • Measurement: Each answer’s top recommendation and brand presence were coded against precompiled per-market brand lists. Matching was case-sensitive to avoid false hits from ordinary words that are also supplier names (notably Turkish).
  • Stability probe: Six identical independent runs per prompt per cell; instability measured as changes in the top recommendation across these repeated runs.
  • API audit: Web_search enabled vs disabled; recorded citations, distinct hosts, utm tags.
  • Limitations noted by authors: six runs per cell is a lower-bound sample; single model family in the API and unspecified model in logged-out UI; one prompt per category; limited time-of-day control.

Implications for AI Economics

  • Measurement validity: Visibility or “search-audit” studies must report query language and egress (exit IP). Measuring visibility in English systematically misrepresents what buyers see in local markets and can entirely miss a firm’s local competitors.
  • Competition and market discovery: Generative interfaces can effectively partition markets by language and egress, changing which firms are discoverable to users. Firms that rely on English-facing optimisation may be invisible to local-language customers; conversely, diaspora users can receive local-market recommendations in their own language.
  • Platform gatekeeping and localisation economics: Language acts as a gating signal (not merely a preference). Platforms’ internal design choices (retrieval, instruction tuning, safety/template policies) can induce market-level exposure effects—potentially amplifying global brands in English and shielding or privileging local incumbents in local languages.
  • Policy and regulatory considerations: The negative-control finding points to the role of category-specific regulation (e.g., VAT, e-invoicing) as a plausible trigger for localisation behaviour. Regulators and competition authorities should consider how algorithmic localisation affects market access, especially for regulated services where legal compliance requires local solutions.
  • Measurement practice recommendations: (1) Always specify query language and egress IP; (2) sample repeatedly because single-run evidence is noisy; (3) if using APIs, report whether external retrieval/web_search was enabled and note that API calls may not reflect end-user geography; (4) include multiple languages and exit locations to detect separable effects of language vs location.
  • Research directions: Identify the mechanism (retrieval vs instruction-tuning vs policy templates), expand categories to disentangle supply-side availability from regulation-driven localisation, and study economic impacts on entry, advertising strategy, and multilingual marketing.

Limitations to carry forward: small per-cell sample size, single model family, single prompt per category, and inability in this study to conclusively identify the internal mechanism (retrieval vs tuning vs list-construction rules).

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper reports a carefully controlled, replicated measurement exercise showing large, consistent differences across languages and exit IPs; however, the per-cell sample size is small (six runs), only a small set of countries, categories, and prompts were used, the web UI model is undisclosed, causal mechanism (retrieval vs. instruction tuning vs. safety filters) is not identified, and results are bounded to a short time window. Methods Rigormedium — Design strengths include independent manipulation of language and egress, logged-out contexts, negative control, case-sensitive coding corrections, and API vs UI comparison; weaknesses include small n per cell (6), few prompts per category, limited temporal coverage, undisclosed production model behind the web UI, potential residual measurement artefacts (e.g., brand names that are common words), and inability to inspect retrieval citations for UI cells. Sample234 usable runs collected across 11 experimental cells on 29–30 Aug 2026: logged-out ChatGPT web UI with residential/ISP exit IPs in Berlin (DE), Oslo (NO), Istanbul (TR) and Tallinn (EE) and an OpenAI API arm (gpt-5.6-terra) with web_search enabled/disabled; six identical runs per prompt per cell; treatment category = accounting software for freelancers (primary), control category = project management; multiple query languages (local languages, English, and Russian in Tallinn); brand presence coded against pre-assembled per-market lists. Themesadoption innovation IdentificationControlled factorial probe that independently manipulates query language and exit IP (egress), uses fresh logged-out browser contexts and an API arm (with web_search on/off), six identical runs per cell, per-market brand lists coded case-sensitively, and a negative control category to distinguish market supply from regulatory effects; comparison across replicated cells and Fisher exact tests for key contrasts. GeneralizabilitySmall per-cell sample (six runs) limits confidence in stability estimates and rare-event detection, Only four geographic egresses and a few languages—results may not hold in other countries/languages, Single short collection window (two days) — model behaviour may change with model updates or time, Web UI model undisclosed; API pinned to one model family—findings may differ for other models or versions, Only a few product categories (accounting, project management, and a few others) — category-specific regulation may drive effects, Prompts limited (one prompt per category generally), so prompt phrasing sensitivity not fully explored, Cannot inspect/UI citations in many cells, so underlying mechanism (retrieval vs. tuning vs. safety lists) is not identified

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The top recommendation changed across four of six prompts in the Berlin web-interface arm, four of six in the Oslo web-interface arm, and four of six in both API conditions, regardless of whether web search was enabled. Decision Quality negative Run-to-run stability of the top commercial recommendation
Reading fidelity high
Study strength medium
n=234
4 of 6 prompts were unstable in each reported condition
0.48
Query language strongly affects whether local suppliers appear in commercial recommendations: across four official-language cells, a global supplier appeared in only 1 of 24 runs, whereas local suppliers appeared in 0 of 12 English-language runs from Istanbul and Tallinn. Market Structure positive Presence of local versus global suppliers in recommendations
Reading fidelity high
Study strength medium
n=36
Global supplier in 1 of 24 official-language runs; local supplier in 0 of 12 English-language runs
0.48
On the same exit connection, asking in the local language produced substantially more local accounting-software suppliers than asking in English. Market Structure positive Local supplier recommendation or winner status
Reading fidelity high
Study strength medium
n=24
Fisher's exact test p = 0.0152 for Fiken in Norway; one-sided p = 0.0076 and 0.0303 for Germany; combined p = 0.0022
0.48
Holding the query language fixed in Turkish and changing only the exit IP moved the recommended supplier set from Turkish suppliers to German suppliers, while the answer remained in Turkish. Market Structure positive Country of recommended suppliers and answer language
Reading fidelity high
Study strength medium
n=12
From Berlin: German suppliers Lexware and sevdesk appeared in 5 of 6 runs; from Istanbul: Turkish suppliers appeared in all six runs
0.48
Russian, a minority language on the Estonian connection, produced an intermediate localization pattern: Estonian suppliers appeared in 4 of 6 runs, compared with 0 of 6 in English and local suppliers appearing in every Estonian-language run. Market Structure mixed Presence of Estonian and global suppliers by query language
Reading fidelity high
Study strength medium
n=18
Estonian suppliers in 4/6 Russian runs, 0/6 English runs, and at least one local supplier in 6/6 Estonian runs
0.48
Knowing the user's country did not by itself cause the interface to recommend local suppliers: every Estonian- and Russian-language run named Estonia in the answer text, but local Estonian suppliers appeared in only four of six Russian runs and none of six English runs. Decision Quality negative Local supplier recommendation conditional on country awareness
Reading fidelity high
Study strength medium
n=18
Country named in 6/6 Estonian runs and 6/6 Russian runs; local supplier named in 4/6 Russian runs and 0/6 English runs
0.48
The project-management negative control showed no language effect: no domestic supplier appeared in any language or exit-country cell, while global suppliers dominated the recommendations. Market Structure null_result Domestic supplier presence in project-management recommendations
Reading fidelity high
Study strength medium
n=42
0 domestic suppliers in all 42 control runs
0.48
The API returned citations when web search was enabled but none when web search was disabled. Regulatory Compliance positive Citation presence and number of distinct citation hosts
Reading fidelity high
Study strength high
n=72
200 citations across 36 search-enabled runs; 0 citations across 36 search-disabled runs
0.8
The paper argues that the observed language and location effects are better characterized as a market-selection gate than as a simple preference for particular brands, but it does not identify the mechanism implementing the gate. Market Structure mixed Mechanism of market localization in commercial recommendations
Reading fidelity high
Study strength low
n=6
0.24

Notes