1 cumulative citations
View corpus contextLarge-scale AI inference is emerging as critical "cognitive infrastructure" that incumbents can use to foreclose rivals through non-price frictions like latency, routing, and feature gating; the author proposes a narrowly targeted, auditable neutrality regime—quality-of-service parity, routing transparency, and FRAND-style non-discrimination—enforced only when observable evidence shows functional gatekeeper power.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
As generative AI commercializes, competitive advantage is shifting from one-time model training toward continuous inference, distribution, and routing. At the frontier, large-scale inference can function as cognitive infrastructure: a bottleneck input that downstream applications rely on to compete, controlled by firms that often compete downstream through integrated assistants, productivity suites, and developer tooling. Foreclosure risk is not limited to price. It can be executed through non-price discrimination (latency, throughput, error rates, context limits, feature gating) and, where models select tools and services, through steering and default routing that is difficult to observe and harder to litigate. This essay makes three moves. First, it defines cognitive infrastructure as a falsifiable concept built around measurable reliance, vertical incentives, and discrimination capacity, without assuming a clean market definition. Second, it frames theories of harm using raising-rivals'-costs logic for vertically related and platform markets, where foreclosure can be profitable without anticompetitive pricing. Third, it proposes Neutral Inference: a targeted, auditable conduct approach built around (i) quality-of-service parity, (ii) routing transparency, and (iii) FRAND-style non-discrimination for similarly situated buyers, applied only when observable evidence indicates functional gatekeeper status.
Summary
Main Finding
The paper argues that in the era of commercial generative AI, inference (serving/model access, routing, and related API features) can become a “cognitive infrastructure” — a bottleneck input controlled by firms that also compete downstream. When that occurs, anticompetitive foreclosure will often be effected via non-price discrimination (latency, throughput, errors, context limits, feature gating, steering/routing) rather than by headline pricing. The authors propose a targeted, auditable conduct framework called Neutral Inference (QoS parity, routing transparency, FRAND-style access) that should be triggered only when empirical evidence indicates functional gatekeeper status.
Key Points
- Inference-as-competitive-surface: Commercial differentiation and monetization occur at inference and delivery (API tiers, latency, context, tool-use), not only at model weights/training.
- Cognitive infrastructure (falsifiable definition): three required conditions—
A) material reliance with costly substitution (switching causes reengineering or quality/cost loss);
B) upstream control with vertical incentives (provider competes downstream);
C) discrimination capacity (non-price levers to raise rivals’ costs or degrade quality). - Theory of harm: Foreclosure is likely implemented through raising-rivals’-costs (RRC) and platform steering rather than price cuts. Keeping rivals “living but not thriving” can be profit-maximizing for integrated gatekeepers.
- Mechanisms of discrimination: (i) QoS differentials (latency, p50/p95/p99, timeouts, error rates), (ii) feature gating (context length, tool use, enterprise wrappers), (iii) steering and default routing (opaque selection favoring first-party services).
- Auditable tests and triggers: measurable QoS wedges and routing bias checks; suggested practical thresholds (e.g., sustained >15% deltas in p95/p99 or denial rates over a 7-day window) after controlling for region, tier, and request class.
- Neutral Inference obligations: (1) QoS parity for similarly situated buyers after objective controls; (2) routing transparency via auditable logs showing eligible/selected tools and whether commercial ties influenced selection; (3) FRAND-style non-discrimination and fast dispute resolution. Enforcement via periodic audits, standardized test suites, interim measures, and penalties tied to inference revenue.
- Institutional fit: Designed as a conduct-based, narrowly targeted remedy compatible with antitrust enforcement (more auditing/compliance than ex ante regulation), and leveragable with EU DMA/AI Act transparency tools.
Data & Methods
- Nature: conceptual/theoretical paper drawing on industrial organization (raising rivals’ costs), platform economics, and regulatory practice. No new empirical dataset; rather it proposes testable empirical protocols.
- Proposed empirical tests and minimal audit protocol:
- QoS measurement: repeated latency and reliability tests (p50/p95/p99 latency, timeouts, error rates, quota denials) across regions, contracted tiers, and request classes; compare first-party vs third-party, flag persistent residual wedges.
- Routing audit: collect and analyze routing logs that show tool eligibility, selected tool(s), ranking, and whether commercial relationships constrained eligibility/ranking.
- Controls: region/topology, tier, request class, colocation effects, and documented objective criteria (abuse prevention, safety). Distinguish legitimate architecture (colocation, caching) from discriminatory degradation via controlled experiments.
- Practical thresholds: example proposed >15% difference in high-percentile latency/denial rates sustained over a rolling 7-day window as a flag for investigation.
- Feasibility claim: Auditing can be done without disclosure of model weights or private user data; relies on observable service-level measurements and logs.
Implications for AI Economics
- Market-power channel: Concentration in training need not be the sole concern; control over inference infrastructure can create durable upstream bottlenecks that shape downstream competition.
- Non-price competition matters: Traditional price-focused antitrust tests can miss durable harms implemented through quality/routing discrimination. Empirical work and policy must broaden to QoS and routing metrics.
- Measurement and monitoring: Economists must develop standardized metrics (p50/p95/p99 latency, error/denial rates), experimental protocols to control for colocation/architecture, and audit-friendly routing measures. These become central inputs for competition analysis.
- Strategic responses by firms: Expect incentives for colocation, caching, feature-wrapping, and opaque routing to be used strategically; also an equilibrium where incumbents keep third parties viable but constrained (“living but not thriving”).
- Policy design: A narrowly targeted, evidence-triggered neutrality duty (Neutral Inference) can be compatible with antitrust institutions while addressing hard-to-observe non-price harms. This suggests regulators and researchers should prioritize development of operational audit tools, standard disclosure formats, and enforcement mechanisms tied to inference revenues.
- Trade-offs for innovation: Imposing QoS parity and routing transparency may reduce some vertical integration efficiencies (e.g., caching benefits) but can preserve downstream contestability. Empirical evaluation is needed to quantify innovation vs contestability trade-offs.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Competitive advantage is shifting from one-time model training toward continuous inference, distribution, and routing. Market Structure | mixed | locus of competitive advantage (shift from training to inference/distribution/routing) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| At the frontier, large-scale inference can function as cognitive infrastructure: a bottleneck input that downstream applications rely on to compete, controlled by firms that often compete downstream through integrated assistants, productivity suites, and developer tooling. Market Structure | negative | degree to which inference acts as a bottleneck (control over a critical input for downstream competition) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Foreclosure risk is not limited to price. It can be executed through non-price discrimination (latency, throughput, error rates, context limits, feature gating) and, where models select tools and services, through steering and default routing that is difficult to observe and harder to litigate. Market Structure | negative | potential for foreclosure via non-price discrimination and steering/default routing |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Cognitive infrastructure can be defined as a falsifiable concept built around measurable reliance, vertical incentives, and discrimination capacity, without assuming a clean market definition. Governance And Regulation | mixed | criteria for labeling an input as 'cognitive infrastructure' (measurable reliance, vertical incentives, discrimination capacity) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Theories of harm can be framed using raising-rivals'-costs logic for vertically related and platform markets, where foreclosure can be profitable without anticompetitive pricing. Market Structure | negative | possibility that foreclosure increases rivals' costs and is profitable absent price-based exclusion |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Neutral Inference: a targeted, auditable conduct approach is appropriate, built around (i) quality-of-service parity, (ii) routing transparency, and (iii) FRAND-style non-discrimination for similarly situated buyers, applied only when observable evidence indicates functional gatekeeper status. Governance And Regulation | positive | implementation of an auditable, non-discriminatory regulatory regime ('Neutral Inference') to mitigate conduct-based foreclosure |
Reading fidelity
high
Study strength
speculative
|
not reported
|