The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI shopping agents still show position bias, but it is weaker and often 'lost-in-the-middle'; crucially, inspection biases rarely change bookings, which overwhelmingly gravitate to a single undominated hotel, implying SEO placement matters less when consumers delegate purchases to AI.

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
Davood Wadi, Yu Ma · August 24, 2026
arxiv rct medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Davood Wadi unresolved corpus identity
  2. Yu Ma unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Davood Wadi provider ID
  2. Yu Ma unresolved corpus identity
In randomized experiments, LLM-based shopping agents inspect many more listings than humans and exhibit a weak, often U-shaped positional inspection bias, but this inspection bias rarely alters final bookings which concentrate on a single undominated listing, and higher reasoning effort reduces positional effects.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.

Summary

Main Finding

AI agents (LLMs) search more deeply than humans and still exhibit position bias when inspecting listings, but that bias is much weaker and non-monotonic (a "lost‑in‑the‑middle" U‑shape). Crucially, the residual position effects on inspection do not reliably translate into different choices: agents overwhelmingly converge on a single undominated listing and always book (100% conversion), meaning placement matters less than the attributes shown on the results page. Increasing an agent’s reasoning effort further attenuates positional bias.

Key Points

  • Depth of search:
    • AI agents inspect substantially more listings per session than humans: means = 1.63 (Claude Sonnet 5), 3.12 (Gemini 3.1 Flash Lite), 4.25 (Gemini 3.7 Flash), 5.83 (Gemini 3.1 Pro) vs human benchmark 1.12.
    • Degenerate consideration sets (one inspection) are rare for some models (0.8%–37.4% vs 93% for humans).
  • Conversion / outside option:
    • Every AI agent booked in 100% of sessions (versus 66% conversion and 34% outside option for humans).
  • Position effect on inspection:
    • Position still predicts inspection causally (order randomized), but effects are 4–10× smaller than for humans. Linear position coefficients on inspection range roughly −0.0002 to −0.0005 (human ≈ −0.0019).
    • For three of four LLMs, inspection probability declines from the top to a minimum around ranks 68–74 then rises toward the bottom (lost‑in‑the‑middle U‑shape); Gemini 3.1 Pro shows mostly primacy without the recency rise.
  • Position effect on choice:
    • Position rarely affects final booking. Choice position coefficients are indistinguishable from zero for two models and very small for others. Average chosen ranks ≈ 44.0–49.7 (neutral = 50.5).
    • Choices concentrate on a single undominated hotel (same modal choice across models), which accounted for 78.2% of all bookings pooled.
  • Reasoning effort:
    • Manipulating reasoning effort (multiple levels across Gemini Flash Lite and Gemini Pro) reduces the lost‑in‑the‑middle curvature; at higher effort levels the positional curvature and the position effect on choice become insignificant.
  • Controls and other predictors:
    • Attributes visible on the results page (review score, price, promotion, chain) strongly predict inspection and choice; review score is a major heuristic driving choices.
  • Heterogeneity:
    • Model behavior differs substantially across LLMs in search depth and inspection patterns; this heterogeneity does not map cleanly to provider or capability tier alone.

Data & Methods

  • Setup:
    • Randomized controlled experiment that randomized the order of 100 hotel listings per session (to eliminate ranking-quality endogeneity).
    • Two-layer information structure: listing page (name, review score, nightly price, total price, aggregated rating) and an inspect tool that returns detailed hotel attributes when called (analogue of a click).
    • Agents received the full 100-item page in the context window and could sequentially call inspect and then submit_choice (or terminate without booking).
  • Models and sample:
    • Main experiment: 4 LLMs (Google Gemini 3.1 Pro, Gemini 3.7 Flash, Gemini 3.1 Flash Lite; Anthropic Claude Sonnet 5), 500 independent sessions per model → 2,000 sessions.
    • Follow-ups: reasoning effort manipulation across 7 cells (Gemini Flash Lite: 4 effort levels; Gemini Pro: 3 levels), 500 replications per cell → 3,500 sessions.
  • Measures and estimation:
    • For each session×hotel: Inspected (1 if inspect tool used), Chosen (1 if booked). Position indexed 1–100 random per session.
    • Controls: price (nightly), review score (aggregated), chain (brand binary), promotion (discount badge binary).
    • Analysis: linear probability models with session-clustered standard errors; tests for nonlinear (quadratic) position terms to identify U‑shape.
  • Benchmarks:
    • Results compared to human field benchmark (Ursu 2018), acknowledging differences in list length and some displayed attributes.

Implications for AI Economics

  • Market value of ranking slots will change under delegated search:
    • The traditional premium for top positions (driven by limited human attention and sequential scanning) weakens when consumers delegate to agents that can process full pages and deliberate.
    • Platforms that monetize placement (paid search, sponsored slots, auctions) may see reduced willingness to pay for top slots; price discrimination strategies and auction design may need revision.
  • Importance of on‑page attributes over placement:
    • For agentic delegation, visible quality signals (review scores, price, promotions) matter more for final demand than rank. Managers should reallocate SEO/marketing spend toward improving and signaling attributes rather than only chasing top placement.
  • Platform and recommender design:
    • Recommender interfaces and APIs should prioritize clear, structured attribute fields that agents can reliably use (since agents rely on attributes), and consider how context length and information ordering interact with model retrieval behaviors (lost‑in‑the‑middle).
    • Providing structured metadata (high-quality ratings, canonical prices) may be more effective than position tweaks for influencing agent-mediated sales.
  • Strategic behavior and competitive dynamics:
    • With agents converging on undominated options, winner-take-most dynamics from a quality signal perspective may intensify (one listing capturing large share), shifting competition toward improving objective dominance (better price/score combinations).
  • Policy and consumer welfare:
    • Agents’ tendency to always book and to focus on a single modal option raises questions about user control, outside-option use, and necessity of transparency/disclosure for delegated decisions.
    • Regulators and platforms may need to consider standards for how search results are structured for agent consumption, and whether agents should be required to surface alternatives or explanations.
  • Research directions:
    • Evaluate real-world outcomes where users actively choose to delegate to personal agents (endogenous delegation), other product categories, multi-stakeholder settings (advertisers/platforms), and strategic listing behavior when agents are common.
    • Study how personalization and agent objectives (e.g., risk aversion, client preferences) interact with position/attribute effects and market equilibria.

Limitations worth noting: lab-style agent experiments use default prompts and provider APIs (sampling parameters often not adjustable); external validity depends on how real users compose prompts, select agents, and constrain agent behavior; and results may vary with different listing lengths, domains, or richer attribute sets.

Assessment

Paper Typerct Evidence Strengthmedium — Strong internal validity for how these specific LLMs behave in the constructed environment (randomization, large N, within-session clustering, robustness checks). External validity is limited: simulated tool calls and prompts may not reflect real-world delegated purchases, only four proprietary LLM families were tested, and the prompt/utility framing (agents always book under many conditions) may drive results. Methods Rigorhigh — Careful experimental design (full randomization of positions), substantial replication (500 sessions per cell), clustered SEs, controls for observable listing attributes, tests for nonlinearity, and a follow-up factorial manipulation of reasoning effort; potential weaknesses are reliance on linear probability models (transparent but limited), limited exploration of prompt sensitivity, and only a handful of LLMs/providers. SampleMain experiment: 2,000 independent LLM sessions (500 each) across four proprietary LLM models: Google Gemini 3.1 Pro, Gemini 3.7 Flash, Gemini 3.1 Flash Lite, and Anthropic Claude Sonnet 5; each session received a fully randomized ordering of 100 hotel listings and the agent could 'inspect' listings via a tool or submit a booking; follow-up experiments: 3,500 sessions manipulating reasoning-effort levels (7 cells × 500). Human benchmark summary statistics from Ursu (2018) are used for comparison. Themeshuman_ai_collab adoption IdentificationRandomized ordering of 100 hotel listings within each agent session (2000 main sessions, plus 3500 in follow-ups) gives causal identification of position effects; additional randomized manipulation of LLM reasoning-effort levels; comparisons to a human benchmark (Ursu 2018). Standard errors clustered at the session level; reduced-form linear probability models and nonlinear (quadratic) checks reported. GeneralizabilitySimulated environment: tool-based 'inspect' and 'submit_choice' abstractions may not match real-world UI, API constraints, or user-agent integrations., Prompt and persona dependence: results may be sensitive to wording, instruction to recommend/book, and to the exact delegation framing., Limited model coverage: four LLM variants from two providers — behavior may differ for other models, future model updates, or open-source systems., Single task/domain: hotel search with a specific itinerary and static listing attributes; other purchase contexts or longer multi-step tasks may produce different dynamics., Agent incentives: agents always booking in many conditions may reflect prompt or default sampling settings rather than general delegated consumer behavior., Temporal dynamics: LLM capabilities, interfaces, and marketplace deployments evolve rapidly, limiting long-run external validity.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI agents inspect substantially more hotel listings per session than human consumers: the four tested models averaged between 1.63 and 5.83 inspections per session, compared with 1.12 for humans. Consumer Welfare positive Number of hotel listings inspected per search session
Reading fidelity high
Study strength medium
n=2000
1.63–5.83 inspections per session versus 1.12 for humans
0.6
All tested AI agents booked a hotel in every session, whereas human consumers in the benchmark booked 66% of the time and selected the outside option in 34% of sessions. Consumer Welfare positive Hotel booking/conversion rate
Reading fidelity high
Study strength medium
n=2000
100.0% conversion for each AI model versus 66.0% for humans
0.6
Listing position negatively predicts whether an AI agent inspects a hotel, but the position effect is four to ten times weaker than for human consumers. Consumer Welfare negative Probability that a displayed hotel is inspected
Reading fidelity high
Study strength high
n=2000
Position coefficient −0.0002 to −0.0005 for AI agents versus −0.0019 for humans
1.0
For most tested LLMs, inspection probability is non-monotonic in position: inspection declines from the top of the page to a minimum around ranks 68–74 and then rises toward the bottom. Consumer Welfare mixed Inspection probability by listing rank
Reading fidelity high
Study strength medium
n=2000
Minimum inspection probability at approximately ranks 68–74
0.6
The effect of listing position on final hotel choice is heterogeneous across models: it is statistically indistinguishable from zero for Gemini Flash and Gemini Pro, but negative and statistically significant for Claude Sonnet and Gemini Flash Lite. Consumer Welfare mixed Probability that a hotel is selected as the final booking
Reading fidelity high
Study strength high
n=2000
Choice-position coefficients from −0.000008 to −0.000080
1.0
The four LLMs strongly converge on the same hotel choice: one undominated listing captured 78.2% of all bookings across the 2,000 main-experiment sessions. Consumer Welfare positive Concentration and consistency of final hotel choices
Reading fidelity high
Study strength medium
n=2000
78.2% of bookings, or 1,563 of 2,000 sessions
0.6
Increasing LLM reasoning effort reduces the position effect on final choice, with the position effect becoming statistically insignificant at the highest tested effort level for both Gemini Flash Lite and Gemini Pro. Consumer Welfare negative Position effect on final booking choice
Reading fidelity high
Study strength medium
n=3500
Position effect not significant at the highest reasoning-effort level
0.6
Changing reasoning effort affects search depth in opposite directions across the two tested models: inspections fell from 3.12 to 1.91 per session for Gemini Flash Lite but rose from 0.39 to 5.83 for Gemini Pro. Consumer Welfare mixed Number of hotel inspections per session
Reading fidelity high
Study strength medium
n=3500
Flash Lite: 3.12 to 1.91 inspections; Pro: 0.39 to 5.83 inspections
0.6

Notes