0 cumulative citations
View corpus contextShared model origins can enable profitable, communication‑free collusion that leaves each bidder’s marginal bids unchanged, rendering standard price‑level audits provably blind — and evidence from language‑model agents and Ethereum relay auctions suggests this is a real regulatory blind spot.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.
Summary
Main Finding
A profitable form of collusion can be implemented by coupling bidders only through the joint distribution (the copula) of their unexplained bid components while leaving each bidder’s own marginal bid law identical to the competitive benchmark. Any audit or statistical test that inspects only a single bidder’s marginal price or bid history (the class of "price-level" tests common in the literature) has power equal to its nominal size α against this conduct — i.e., it is blind by construction, at every sample size and at every coupling strength up to perfect comonotonicity. The paper shows this construction is economically meaningful (it can transfer real rent), appears in real language-model deployments, and is already at damaging levels in a deployed market (Ethereum block building).
Key Points
-
Mechanism (construction)
- Fix each bidder i’s marginal unexplained bid distribution Fi, and couple the unexplained residuals across bidders via a shared latent draw zt refreshed each round.
- No communication channel required: zt can be derived from public round data via a keyed pseudorandom function or arise naturally when multiple deployments share the same base model.
- The comonotone coupling (perfect positive dependence) minimizes the expected maximum bid, so coupling alone lowers seller revenue even though marginals are unchanged.
-
Detection impossibility for price-level audits
- Proposition: any test A that uses only one bidder’s history {(bit, xt)} has PrP(A rejects) = α under the coupled conduct (Lemma 1 construction). Power equals size exactly.
- This is not low power that more data fixes; it is an identity that holds up to comonotonicity because the single-bidder process law is unchanged.
-
Why existing audit ideas fail here
- Price-level tests: blind because marginals unchanged.
- Regret-based audits (calibrated/swap regret on single-agent transcripts): also fail because the honest market itself carries irreducible noise and positive calibrated regret; the coupling does not raise excess regret above that ambient level and can reduce it, so a regulator calibrated to pass the honest market will pass the conspiracy too.
- Single-agent serial tests would detect the conduct if the shared draw is persistent across rounds; the impossibility relies on the shared draw being refreshed independently every round.
-
Economic significance
- Monte Carlo and market-calibration show the conduct can extract seller-side rent of comparable order to previously reported supracompetitive price effects.
- When outsiders and dispersion are present, coupling (high ρ) raises coalition surplus while lowering seller revenue.
-
Empirical evidence
- Language-model agents: 20 models from 19 independent developers, 3 prompts each, 5 temperatures, 200 rounds per temperature. After regressing out observable order features (auditors of increasing richness), two deployments of the same base model show positive residual correlation (+0.053 under the strongest auditor, clustered 95% CI [0.030, 0.078]) while cross-base pairs do not (≈0.0001). Correlation weakens monotonically with sampling temperature (p = 0.002).
- Ethereum MEV-Boost relay data: 24 days, 77,684 bids, 39 bidders. The honest population’s pairwise dependence is so dispersed that to hold a 5% false-positive rate a screen must set a correlation threshold between +0.50 and +0.81 — 20–32× the family-wise sampling threshold — and this floor does not decline with the audit window.
- Concentration / counting: Resolving multi-identity operation into fewer operators increases measured concentration substantially (e.g., resolving 40 identities into 23 operators raises the Herfindahl index by +247.5%; adding behavioral clusters from bid streams reaches +324.5%).
Data & Methods
-
Theoretical model and proofs
- Auction model: repeated first-price auction rounds. Each bid decomposed into public component git and unexplained residual σεit with marginal law Fi; κ = sd(σε)/sd(g) measures unexplained share.
- Lemma 1: constructs a zero-communication coupling using a PRF keyed on public round data (or shared base model outputs) to produce zt; sets εit = F−1 i (Φ(√ρ zt + √(1−ρ) eit)) so each εit marginal = Fi for any ρ ∈ [0,1].
- Propositions: show marginal audit power = α (Proposition 1/3), comonotone coupling minimizes expected max (seller revenue) (Proposition 2), and regret audits cannot reliably separate the conduct because ambient calibrated regret is nonzero and not increased by coupling (Proposition 4).
-
Simulation
- Monte Carlo validates that single-agent tests remain at nominal size α across sample sizes up to N = 20,000 and across coupling strengths; pairwise tests on residuals have high power and scale as typical N−1/2.
-
Language-model experiments
- 60 agent deployments (20 models × 3 prompts), run 200 rounds per temperature across five temperatures.
- Residual correlations computed after regressing out observable order features under three auditors: (A) linear in 2 features, (B) linear in all 6 features, (C) rich model with squares and interactions, cross-fitted out of sample; C is the strictest.
- Correlations clustered by developer; permutation tests show monotone decline with temperature (p = 0.002).
-
Market calibration (Ethereum)
- 24-day MEV-Boost relay dataset, 77,684 bids from 39 bidders.
- Measured distribution of pairwise residual correlations among ostensibly honest bidders to compute the detection floor at fixed false-positive rates; showed floor is large and stable with window length.
-
Concentration / entity-resolution exercise
- Re-assigning identities to operators (entity resolution and behavioral clustering) and computing Herfindahl index changes to show sensitivity of market power measures to counting rules.
Implications for AI Economics
-
Auditing and regulation
- Price-level audits (and many existing empirical frameworks that ask "are prices supracompetitive?") are insufficient by design to detect collusion implemented via joint residual coupling while marginals are unchanged.
- Regulators should not rely solely on single-agent marginal tests or regret-calibrated single-agent audits; to detect this mechanism audits must examine joint behavior across agents (pairwise or higher-order dependence) or intervene on other levers.
- Because lawful multi-identity operation and conspiratorial coupling are behaviorally indistinguishable here, a tractable regulatory target is counting — i.e., accurate entity resolution and mapping identities to independent decision-makers. Measured concentration is extremely sensitive to counting rules.
-
Practical mitigation levers
- Sampling temperature (randomisation in agent outputs) is a promising deployment parameter: higher temperature reduces coupling empirically. Regulators might consider minimum randomness requirements or standards for deployment diversity to reduce extractable rent from shared-model coupling.
- Require disclosure or provenance audits of shared model sources, or mandate diversity-of-sources in critical market decisions, to reduce chances that multiple market actors share the same model/computation producing correlated residuals.
-
Rethinking "collusion" in AI-mediated markets
- The classic antitrust focus on meetings-of-the-minds or explicit communication misses a class of economically harmful dependence that arises from shared models or common upstream functions.
- Third parties (model providers) can capture rent by fixing the copula; policy should consider the role of upstream model standardisation in market power.
- Empirical work on algorithmic collusion should include joint-dependence screens (pairwise residual correlations, copula tests) and calibrate detection thresholds to the ambient distribution of honest dependence.
-
Research directions
- Develop statistically powerful joint tests for copula dependence that account for the wide ambient variation in honest pairwise dependence (i.e., control false positives while retaining power).
- Scale-out randomized-deployment experiments to quantify how deployment knobs (temperature, prompt randomization, model ensembling) affect extractable rent.
- Build robust entity-resolution methods tailored to auction/bidding streams and evaluate their use as regulatory primitives (e.g., in Herfindahl or other concentration measures).
Limitations and caveats - The impossibility result depends critically on the shared draw being refreshed independently every round; persistent common signals would create detectable single-agent serial structure. - The conduct is profitable in regimes with nontrivial unexplained bid dispersion; if outsiders or the environment have very low unexplained variance, coupling may not benefit the coalition. - Detection by examining joint behavior requires data access across participants that regulators or auditors may not have; the policy shift to counting and provenance entails new data-collection and legal/institutional measures.
Summary takeaway Shared-model or shared-function coupling of agents can implement economically meaningful, zero-communication collusion that leaves each agent’s marginal bids identical to competitive benchmarks, making standard price-level and single-agent regret audits blind by construction. Regulators and empirical researchers should shift attention from marginal price tests to joint-dependence analysis, entity resolution (counting), and deployment rules (e.g., mandated randomness or provenance disclosure) to mitigate the risk.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| A conspiracy can preserve every bidder's marginal bid distribution at the competitive level while changing only the dependence structure between bidders. Task Allocation | mixed | Bid distribution and cross-bid dependence |
Reading fidelity
high
Study strength
high
|
not reported
|
| Any audit based only on one bidder's price or bid history has power exactly equal to its false-positive rate against the fixed-marginal conspiracy, for every coupling strength up to perfect coupling. Governance And Regulation | null_result | Detection-test power |
Reading fidelity
high
Study strength
high
|
n=20000
power = α
|
| Coupling bidders' unexplained bid components lowers the seller's expected revenue while leaving the bidders' marginal bid distributions unchanged. Consumer Welfare | negative | Seller revenue, represented by the expected winning bid |
Reading fidelity
high
Study strength
high
|
−52% expected maximum residual bid
|
| In the authors' simulations, pairwise tests detect the coupling with power approaching one, whereas single-agent tests remain at nominal size regardless of coupling strength and sample size. Governance And Regulation | mixed | Statistical detection power |
Reading fidelity
high
Study strength
medium
|
n=20000
pairwise-test power = 1; single-agent power = 0.05
|
| Residual bids from two deployments of the same language model are positively correlated, while residual bids from different models are approximately uncorrelated after controlling for observable order features. Ai Safety And Ethics | positive | Residual bid correlation |
Reading fidelity
high
Study strength
medium
|
n=20
+0.053 correlation versus +0.0001 across models
|
| Increasing sampling temperature monotonically weakens residual bid coupling between deployments of the same model. Ai Safety And Ethics | negative | Residual bid correlation as a function of sampling temperature |
Reading fidelity
high
Study strength
medium
|
n=20
correlation declined from +0.070 to −0.004 under auditor (C); p = 0.002
|
| The honest population of Ethereum MEV-Boost bidder pairs has substantial dependence, creating a positive detection threshold that does not disappear as the audit window grows. Governance And Regulation | negative | Minimum detectable residual-dependence threshold |
Reading fidelity
high
Study strength
medium
|
n=77684
dependence floor of +0.50 to +0.81
|
| Under the paper's simulated market conditions, a five-member coupled coalition can capture a substantial share of seller revenue while reducing seller revenue. Consumer Welfare | negative | Seller revenue |
Reading fidelity
high
Study strength
medium
|
4.6% to 10.6% of seller revenue
|
| Regret-based audits also fail to distinguish the fixed-marginal conspiracy from an honest noisy market under the paper's conditional assumption that unexplained bid dispersion represents off-best-response noise. Governance And Regulation | null_result | Excess calibrated swap regret used for audit classification |
Reading fidelity
high
Study strength
low
|
R(ρ) ≤ Rcomp and R(ρ) decreases with ρ
|
| Resolving multiple bidding identities into underlying operators substantially increases measured market concentration. Market Structure | positive | Herfindahl concentration index |
Reading fidelity
high
Study strength
medium
|
n=40
+247.5% from entity resolution; +324.5% with behavioral clustering
|