Generative search favors local suppliers only when users ask in the local language: English queries on the same connection return global brands while exit IP determines which national market is referenced; the browser UI and API are equally unstable, so single snapshots can be misleading.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third, language and location are separable and act on different things: holding the query language fixed and moving only the exit IP moves the market whose brands are named while the answer stays in the query language. We show this on two unrelated pairs, Turkish asked from Berlin and Russian asked from Tallinn, and in both the answer names the resident country's suppliers. A minority language occupies a middle tier: Russian asked from Estonia names an Estonian supplier in 4 of 6 runs and a global one in all six, where Estonian names a local supplier in every run and English names none. A negative control in a second category, coded with the same instrument, shows no language effect at all, and disconfirms our own expectation: that category does have domestic suppliers and none was named in any language, which points the explanation at whether a category is nationally regulated rather than at whether it is nationally supplied.
Summary
Main Finding
A generative search interface’s commercial recommendations are strongly gated by the query language (which decides whether localisation is attempted) and separately by the exit IP (which decides which national market is treated as the user’s). Asked in a local language, the model overwhelmingly names local suppliers; asked in English from the same connection it names global suppliers. The top recommendation is also non-deterministic: identical runs often return different top picks.
Key Points
- Language as gate: When the query language matched the country’s official language, local suppliers dominated (e.g., Estonian-language queries from Tallinn named Estonian accounting vendors in 6/6 runs). The same connection asked in English named global vendors (Estonia: local suppliers 6/6 in Estonian vs 0/6 in English).
- Exit-IP market selection is separable from language: holding query language fixed and changing only the exit IP changed the market named while the answer stayed in the same language (example: Turkish queries from Istanbul → Turkish suppliers; Turkish queries from Berlin → German suppliers, still in Turkish).
- Three-tier ordering on one connection: On the Tallinn connection, Estonian (official) > Russian (local minority) > English (global) in terms of producing local supplier mentions. Russian often produced a mix of local + global; English produced only global names.
- Instability: Across six identical runs per prompt, the top recommendation changed on 4/6 prompts in both the logged-out web UI and in the API (with and without web search). Instability rates were equal across surfaces, implying system-level nondeterminism.
- API citation behavior: API calls with web_search enabled returned citations (200 citations across 36 runs; ~5.6 per answer), while API calls with web_search disabled returned none. The API also lacked user geography; it returned U.S.-centric answers when asked in English unless exit IP was otherwise specified by the collector.
- Negative control: For a second category (project-management software), domestic suppliers were not named in any language or exit-IP. This disconfirms the simple “exists-a-domestic-supplier → language effect” expectation and suggests category-specific factors (likely regulatory constraints) determine whether localisation is invoked.
- Statistical support: Two independent replications (Germany and Norway) show a significant language effect for local-vs-English winners (combined Fisher’s method p ≈ 0.0022).
Data & Methods
- Scope: 234 usable runs collected 29–30 Aug 2026 across 11 cells (combinations of surface, exit IP and query language); six identical runs per prompt per cell (plus some discarded/recollected runs for instrument checks).
- Cells: exits included Berlin (DE), Oslo (NO), Istanbul (TR) and a residential Tallinn (EE) connection. Surface: logged-out ChatGPT web interface and OpenAI API (gpt-5.6-terra) with web_search on/off.
- Prompts: Treatment = accounting software for freelancers (chosen because each country has domestic suppliers). Control = project management for a small marketing agency.
- Measurement: Each answer’s top recommendation and brand presence were coded against precompiled per-market brand lists. Matching was case-sensitive to avoid false hits from ordinary words that are also supplier names (notably Turkish).
- Stability probe: Six identical independent runs per prompt per cell; instability measured as changes in the top recommendation across these repeated runs.
- API audit: Web_search enabled vs disabled; recorded citations, distinct hosts, utm tags.
- Limitations noted by authors: six runs per cell is a lower-bound sample; single model family in the API and unspecified model in logged-out UI; one prompt per category; limited time-of-day control.
Implications for AI Economics
- Measurement validity: Visibility or “search-audit” studies must report query language and egress (exit IP). Measuring visibility in English systematically misrepresents what buyers see in local markets and can entirely miss a firm’s local competitors.
- Competition and market discovery: Generative interfaces can effectively partition markets by language and egress, changing which firms are discoverable to users. Firms that rely on English-facing optimisation may be invisible to local-language customers; conversely, diaspora users can receive local-market recommendations in their own language.
- Platform gatekeeping and localisation economics: Language acts as a gating signal (not merely a preference). Platforms’ internal design choices (retrieval, instruction tuning, safety/template policies) can induce market-level exposure effects—potentially amplifying global brands in English and shielding or privileging local incumbents in local languages.
- Policy and regulatory considerations: The negative-control finding points to the role of category-specific regulation (e.g., VAT, e-invoicing) as a plausible trigger for localisation behaviour. Regulators and competition authorities should consider how algorithmic localisation affects market access, especially for regulated services where legal compliance requires local solutions.
- Measurement practice recommendations: (1) Always specify query language and egress IP; (2) sample repeatedly because single-run evidence is noisy; (3) if using APIs, report whether external retrieval/web_search was enabled and note that API calls may not reflect end-user geography; (4) include multiple languages and exit locations to detect separable effects of language vs location.
- Research directions: Identify the mechanism (retrieval vs instruction-tuning vs policy templates), expand categories to disentangle supply-side availability from regulation-driven localisation, and study economic impacts on entry, advertising strategy, and multilingual marketing.
Limitations to carry forward: small per-cell sample size, single model family, single prompt per category, and inability in this study to conclusively identify the internal mechanism (retrieval vs tuning vs list-construction rules).
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The top recommendation changed across four of six prompts in the Berlin web-interface arm, four of six in the Oslo web-interface arm, and four of six in both API conditions, regardless of whether web search was enabled. Decision Quality | negative | Run-to-run stability of the top commercial recommendation |
Reading fidelity
high
Study strength
medium
|
n=234
4 of 6 prompts were unstable in each reported condition
|
| Query language strongly affects whether local suppliers appear in commercial recommendations: across four official-language cells, a global supplier appeared in only 1 of 24 runs, whereas local suppliers appeared in 0 of 12 English-language runs from Istanbul and Tallinn. Market Structure | positive | Presence of local versus global suppliers in recommendations |
Reading fidelity
high
Study strength
medium
|
n=36
Global supplier in 1 of 24 official-language runs; local supplier in 0 of 12 English-language runs
|
| On the same exit connection, asking in the local language produced substantially more local accounting-software suppliers than asking in English. Market Structure | positive | Local supplier recommendation or winner status |
Reading fidelity
high
Study strength
medium
|
n=24
Fisher's exact test p = 0.0152 for Fiken in Norway; one-sided p = 0.0076 and 0.0303 for Germany; combined p = 0.0022
|
| Holding the query language fixed in Turkish and changing only the exit IP moved the recommended supplier set from Turkish suppliers to German suppliers, while the answer remained in Turkish. Market Structure | positive | Country of recommended suppliers and answer language |
Reading fidelity
high
Study strength
medium
|
n=12
From Berlin: German suppliers Lexware and sevdesk appeared in 5 of 6 runs; from Istanbul: Turkish suppliers appeared in all six runs
|
| Russian, a minority language on the Estonian connection, produced an intermediate localization pattern: Estonian suppliers appeared in 4 of 6 runs, compared with 0 of 6 in English and local suppliers appearing in every Estonian-language run. Market Structure | mixed | Presence of Estonian and global suppliers by query language |
Reading fidelity
high
Study strength
medium
|
n=18
Estonian suppliers in 4/6 Russian runs, 0/6 English runs, and at least one local supplier in 6/6 Estonian runs
|
| Knowing the user's country did not by itself cause the interface to recommend local suppliers: every Estonian- and Russian-language run named Estonia in the answer text, but local Estonian suppliers appeared in only four of six Russian runs and none of six English runs. Decision Quality | negative | Local supplier recommendation conditional on country awareness |
Reading fidelity
high
Study strength
medium
|
n=18
Country named in 6/6 Estonian runs and 6/6 Russian runs; local supplier named in 4/6 Russian runs and 0/6 English runs
|
| The project-management negative control showed no language effect: no domestic supplier appeared in any language or exit-country cell, while global suppliers dominated the recommendations. Market Structure | null_result | Domestic supplier presence in project-management recommendations |
Reading fidelity
high
Study strength
medium
|
n=42
0 domestic suppliers in all 42 control runs
|
| The API returned citations when web search was enabled but none when web search was disabled. Regulatory Compliance | positive | Citation presence and number of distinct citation hosts |
Reading fidelity
high
Study strength
high
|
n=72
200 citations across 36 search-enabled runs; 0 citations across 36 search-disabled runs
|
| The paper argues that the observed language and location effects are better characterized as a market-selection gate than as a simple preference for particular brands, but it does not identify the mechanism implementing the gate. Market Structure | mixed | Mechanism of market localization in commercial recommendations |
Reading fidelity
high
Study strength
low
|
n=6
|