When shippers delegate carrier choice to LLMs, agents herd on the same few carriers—one carrier can attract up to 76% of requests—creating risky concentration; disclosing each carrier's remaining daily capacity cuts concentration by a third and doubles shipper surplus, while other platform tweaks have little detectable effect.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching market, and which platform design choices contain it. We carried out agent-based simulations in which fifty shipper agents, built on commercial LLMs from OpenAI (GPT), Anthropic (Claude), and Google (Gemini), procure truckload capacity for thirty days. The market implements the rules of digital freight matching: each load is offered down the shipper's ranked list of carriers (waterfall tendering), carriers have daily capacity limits, spot prices respond to congestion, and carrier ratings accumulate with transactions. We found three risks and one remedy that works. Agents converged at once: for a fixed sampled carrier population, the same carrier was the modal first choice of every model on day one, attracting up to 76% of requests. Because each agent picks from its own randomly drawn list of displayed candidates, the platform controls how many options each shipper sees; concentration rose steeply once lists exceeded about ten carriers, with the onset differing across models. Which carriers ended up dominant varied widely from one sampled market to another, and displaying true quality instead of estimated ratings changed neither the level nor this variability (by design, quality affects only what agents see, never delivery outcomes). Against these risks, disclosing each carrier's remaining daily capacity cut concentration by a third and doubled shipper surplus, while vendor diversification, list-order randomization, and popularity display showed no clearly detectable effect. Platform information design, ahead of model choice or model regulation, is the lever that works.
Summary
Main Finding
When firms delegate carrier selection to LLM-based shipper agents, choices concentrate onto a single carrier immediately (algorithmic monoculture). Market feedback—congestion-driven price increases and capacity limits—partially counters that lock-in by pushing demand to unrated entrants, but platform information design is the decisive lever: disclosing carriers’ remaining daily capacity substantially reduces concentration (≈1/3 reduction) and doubles shipper surplus, while other interventions (vendor diversification, list-order randomization, popularity display, or revealing true reliability) do not.
Key Points
- Immediate monoculture: LLM shipper agents (GPT, Claude, Gemini) converge on the same modal first-choice carrier on day one for a fixed carrier draw; a single carrier attracted up to 76% of requests in some runs (e.g., GPT sent 34–38 of 50 daily requests to the same carrier).
- Two feedback loops shape outcomes:
- Rating loop (+): early favorites accrue reviews and reputational advantage, which entrenches them.
- Price/congestion loop (−): when a carrier becomes crowded its price rises, causing price-sensitive agents to trial cheaper, unrated entrants.
- Exposure effect: concentration rises sharply as the number of displayed candidate carriers L increases; the steep increase occurs once L crosses roughly 10→15 for two vendors (Anthropic’s Claude was an exception and remained flat).
- Market variability: which carrier becomes dominant varies widely across independently sampled carrier populations. Showing true underlying carrier quality instead of estimated ratings did not reduce either concentration levels or this variability.
- Cold-start avoided: contrary to standard reputation-cold-start concerns, newcomers receive trials because congestion-induced high prices at the favorite drive agents to try unrated entrants; entrants’ ratings converge toward true quality over 30 days.
- Interventions tested:
- Capacity disclosure (real-time remaining daily capacity): effective — reduced concentration by ~33% and doubled shipper surplus.
- Vendor diversification (splitting shippers across LLMs): no clear effect.
- List-order randomization: no clear effect.
- Popularity display (showing how often carriers are chosen): no clear effect.
- Replacing live estimated ratings with true reliability disclosure: no meaningful effect on concentration or variability.
- Baselines: a deterministic common-score rule concentrates strongly even at small L; pure random choice gives a low-concentration floor. Candidate-set overlap contributes importantly to the exposure-driven concentration.
Data & Methods
- Agent-based simulation embedding real LLM agents:
- Agents: 50 shipper agents per run, implemented with commercial LLMs (OpenAI GPT, Anthropic Claude, Google Gemini) plus algorithmic baselines (deterministic scoring, random choice).
- Carriers: 20 carriers per sampled market; each carrier has initial unit price, hidden true reliability ρ_j, capacity weight, and route specialty drawn independently from [0.3, 0.9]. Initial track records n0 follow a heavy-tailed Pareto-based draw (some incumbents have many reviews; some have none).
- Time horizon: 30 simulated days; each shipper issues one load per day (random origin–destination distance index and weight); waterfall tendering is used (top-3 ranked carriers per shipper — the platform tenders down the list until a carrier with capacity accepts).
- Exposure L: number of candidate carriers displayed to each agent, varied across {3, 5, 10, 15, 20}.
- Trust signals: three conditions — (i) true reliability disclosed directly, (ii) frozen static initial rating, (iii) live endogenous customer ratings.
- Interventions: capacity disclosure (remaining daily capacity shown), list-order randomization, popularity display, vendor-mix (diversify LLMs across shippers).
- Pricing dynamics: carriers update unit prices daily in response to excess-demand pressure; price growth parameter α = 0.10; bounds prevent prices below 1.05×cost and above 1.5 (index).
- Ratings: after each served shipment, a binary review is left (positive with prob ρ_j); displayed rating is running mean and count.
- Experimental scope and logging:
- 226 experimental cells, 5–10 independent simulation runs per condition, ≈190,000 individual carrier choices recorded.
- Every prompt and response was logged for auditability.
- Key outcome measures:
- Concentration metric κ (market concentration over carriers; reported as time series and end states).
- Shipper surplus, carrier revenues, and variability across runs/seeds.
Implications for AI Economics
- Platform design outranks model choice: before focusing on which LLMs firms deploy or regulating models, platform-level information design (notably capacity visibility) is the most effective instrument to prevent harmful concentration and improve welfare in AI-mediated matching markets.
- Algorithmic monoculture is rapid but not irrevocable: correlated model preferences create extreme first-day concentration, but endogenous market responses (congestion pricing, capacity constraints) can enable market repair and give entrants opportunities—so static analyses of monoculture can overstate permanent lock-in.
- Exposure control is critical: platforms control candidate exposure L; keeping L below the transition region (~10) can prevent the steep rise in concentration. The exact threshold and sensitivity vary by model population and candidate-set overlap.
- Simple transparency (revealing true quality) is insufficient: disclosing true reliabilities did not lower concentration or its unpredictability, implying that information revealing intrinsic quality alone cannot counteract correlated agent choice and congestion dynamics.
- Policy and antitrust concerns: LLM-mediated coordination can produce large demand spikes to single firms rapidly, creating operational strain and potential anti-competitive concentration. Regulators and platforms should consider rules or design changes (e.g., capacity signals, application limits) that manage congestion and exposure.
- Research directions: field validation, testing deeper waterfalls (beyond top-3 tenders), richer review processes, adaptive LLMs that learn from market feedback, robustness to model updates, and strategic manipulation risks (e.g., platforms or carriers influencing displayed candidate lists or capacity signals).
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| For a fixed sampled carrier population, the same carrier was the modal first choice of every model on day one, attracting up to 76% of requests. Market Structure | negative | share of first-choice requests per carrier (modal share) |
Reading fidelity
high
Study strength
medium
|
n=50
up to 76% of requests
|
| Concentration rose steeply once lists exceeded about ten carriers, with the onset differing across models. Market Structure | negative | market concentration as a function of displayed list length |
Reading fidelity
high
Study strength
medium
|
n=50
|
| Which carriers ended up dominant varied widely from one sampled market to another. Market Structure | mixed | variation in identity of dominant carriers across market samples |
Reading fidelity
high
Study strength
medium
|
varied widely
|
| Displaying true quality instead of estimated ratings changed neither the level nor this variability (by design, quality affects only what agents see, never delivery outcomes). Market Structure | null_result | market concentration level and variability under true quality vs estimated ratings display |
Reading fidelity
high
Study strength
medium
|
n=50
|
| Disclosing each carrier's remaining daily capacity cut concentration by a third. Market Structure | positive | market concentration (reduction) following disclosure of remaining daily capacity |
Reading fidelity
high
Study strength
medium
|
n=50
cut concentration by a third
|
| Disclosing each carrier's remaining daily capacity doubled shipper surplus. Consumer Welfare | positive | shipper surplus (buyer welfare) after disclosing carriers' remaining daily capacity |
Reading fidelity
high
Study strength
medium
|
n=50
doubled shipper surplus
|
| Vendor diversification, list-order randomization, and popularity display showed no clearly detectable effect. Market Structure | null_result | market concentration and shipper surplus under alternative platform interventions |
Reading fidelity
high
Study strength
medium
|
n=50
|
| Platform information design, ahead of model choice or model regulation, is the lever that works. Governance And Regulation | positive | effectiveness of platform information design versus other interventions in reducing concentration and improving shipper surplus |
Reading fidelity
medium
Study strength
medium
|
n=50
|
| We carried out agent-based simulations in which fifty shipper agents, built on commercial LLMs (OpenAI GPT, Anthropic Claude, Google Gemini), procure truckload capacity for thirty days. Other | null_result | methodological setup (simulation design) |
Reading fidelity
high
Study strength
high
|
n=50
|
| The simulated market implemented digital freight matching rules: waterfall tendering (loads offered down shipper-ranked lists), carriers with daily capacity limits, spot prices responding to congestion, and carrier ratings accumulating with transactions. Other | null_result | features of the simulated market environment |
Reading fidelity
high
Study strength
high
|
not reported
|