The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Reinforcement‑learning employers on online labor platforms can tacitly collude on wages — but that risk is concentrated when few firms compete and disappears if common training techniques like experience replay are used.

Algorithmic wage setting on online labor platforms
Herbert Dawid, Philipp Harting, Michael Neugart · August 01, 2026 · Labour Economics
openalex theoretical low evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Herbert Dawid provider ID
  2. Philipp Harting provider ID
  3. Michael Neugart provider ID

Semantic Scholar

Latest observation:

  1. H. Dawid provider ID
  2. P. Harting provider ID
  3. M. Neugart provider ID
In simulations of online labor markets, firms that learn posted wages via DQN can develop tacitly collusive wage-setting that lowers wages, particularly when few firms compete and experience replay is not used.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Online digital labor platforms are increasingly important as tools for matching employers and workers. They reduce labor market frictions and provide the opportunity for remote work, but concerns have been raised about whether workers receive adequate incomes. We study how the use of machine-learning based algorithms for the determination of wage offers affects workers’ wages on online labor platforms by setting up a simulation framework and analyzing outcomes under different assumptions about platform and algorithm design. Firms use reinforcement-learning to update posted wages on the platform, and heterogeneous workers send applications based on the posted information. We show that, if firms use a deep Q-network (DQN), a state-of-the-art machine learning algorithm, collusive wages emerge when the number of competitors is small, and no experience replay is used. Our findings are robust across many features of the model, including the design of the online labor platform.

Summary

Main Finding

When firms on online labor platforms use state-of-the-art reinforcement learning (a deep Q-network, DQN) to set posted wages, they can learn tacitly collusive wage-setting behavior that harms workers — but this outcome is conditional. Collusive (non-competitive) wage outcomes emerge robustly in simulation when the number of competing firms is small and when a common DQN training technique (experience replay) is not used. The qualitative result holds across many variants of platform and model design.

Key Points

  • Setting: repeated matching on online digital labor platforms where firms post wages and heterogeneous workers apply based on posted information.
  • Strategic learning: firms independently update posted wages using reinforcement learning (DQN); there is no explicit coordination or communication.
  • Main mechanism: with small numbers of competitors and without experience replay, the learning dynamics converge to tacitly collusive wage policies (i.e., joint outcomes different from the competitive benchmark).
  • Role of experience replay: including experience replay in DQN training disrupts the temporal correlations that facilitate collusive dynamics and prevents the emergence of collusion in the authors’ simulations.
  • Competition intensity matters: as the number of competitors increases, collusive outcomes are less likely; the effect depends on market structure.
  • Robustness: the collusion result is robust to many modeled features, including different platform designs and worker heterogeneity/search behavior.

Data & Methods

  • Approach: agent-based simulation / multi-agent reinforcement learning framework (no empirical field data).
  • Agents:
    • Firms: use reinforcement learning (deep Q-network, DQN) to choose posted wage offers over repeated rounds.
    • Workers: heterogeneous in characteristics and application behavior; they respond to posted wages and other posted information in the matching process.
  • Experiments:
    • Vary number of competing firms.
    • Toggle DQN features (notably experience replay vs. no replay).
    • Vary platform design parameters and worker-side primitives (heterogeneity, search frictions).
    • Measure outcomes such as equilibrium wages, firm profits, worker application patterns, and convergence properties.
  • Diagnostics: analyze learned policies and dynamics to identify when tacit collusion emerges and under what algorithmic/design conditions it fails to appear.

Implications for AI Economics

  • Algorithmic tacit collusion extends beyond product pricing: employer-side use of RL can generate anti-competitive wage outcomes in labor markets.
  • Algorithm design choices matter for market outcomes: low-level ML implementation details (e.g., experience replay) can be policy-relevant levers that either facilitate or prevent collusive dynamics.
  • Regulatory and platform policy:
    • Antitrust and labor regulators should consider algorithmic learning dynamics as potential sources of coordinated (tacit) employer behavior.
    • Platforms could mitigate risk by encouraging or requiring algorithmic features that break temporal correlations (e.g., experience replay, enforced randomization) or by auditing posted-wage dynamics.
    • Increasing market competitiveness (more employers) and transparency of posting practices may reduce the likelihood of collusion.
  • Research directions:
    • Empirical validation on real platforms to detect signatures of algorithmic collusion in posted wages.
    • Study of other algorithmic architectures, information structures, and multi-agent learning protocols.
    • Designing regulation-friendly algorithmic designs and platform rules that preserve market efficiency and worker welfare.

Assessment

Paper Typetheoretical Evidence Strengthlow — Findings come entirely from agent-based / multi-agent RL simulations rather than empirical field or experimental data; results are internally valid within the simulated environment and robust to many modeled variants, but lack real-world validation and are sensitive to model/algorithm choices. Methods Rigormedium — The paper systematically varies key parameters (competition intensity, DQN features, platform primitives) and analyzes learned policies and dynamics, which strengthens causal claims within the model; however, reliance on a single RL architecture (DQN), potential sensitivity to hyperparameters and training regimes, and simplified worker/platform representations limit methodological rigor relative to empirical identification. SampleSynthetic data from agent-based simulations: multiple firm agents using DQN to post wages in repeated rounds on a digital labor platform; heterogeneous simulated workers choose whether and where to apply based on posted information and search frictions; experiments vary firm count, inclusion of DQN experience replay, platform rules, and worker heterogeneity. No real-world or observational data used. Themeslabor_markets governance adoption inequality IdentificationControlled multi-agent simulation experiments: use of deep Q-network (DQN) agents for firms in a repeated matching market, with counterfactual manipulations (number of firms, presence/absence of experience replay, platform design and worker heterogeneity) to compare equilibrium wage and profit outcomes across simulated conditions. GeneralizabilityResults are from simulated environments and may not map directly to real platforms with richer institutional constraints and strategic sophistication., Findings depend on DQN architecture, hyperparameters, and training procedures; other learning algorithms may behave differently., Worker behavior and matching primitives are simplified and may not capture real application/search dynamics or multi-dimensional posted information., Scale effects: behavior with few firms may not generalize to large, fragmented labor markets., Omitted features such as regulation, reputational mechanisms, multi-period contracts, and heterogenous firm objectives could alter outcomes.

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In agent-based simulations of online labor platforms, firms using deep Q-networks to set posted wages can learn tacitly collusive wage-setting behavior without explicit coordination or communication. Market Structure negative Whether firms converge to non-competitive, tacitly collusive wage policies
Reading fidelity high
Study strength low
not reported
0.06
Collusive wage outcomes emerge robustly in the simulations when the number of competing firms is small and experience replay is omitted from DQN training. Market Structure negative Emergence of collusive or non-competitive posted-wage outcomes
Reading fidelity high
Study strength low
not reported
0.06
Including experience replay in DQN training disrupts the temporal correlations associated with collusive dynamics and prevents collusion from emerging in the authors' simulations. Market Structure positive Emergence or prevention of tacit wage collusion
Reading fidelity high
Study strength low
not reported
0.06
Increasing the number of competing firms makes collusive wage outcomes less likely. Market Structure positive Likelihood of tacitly collusive wage-setting behavior
Reading fidelity high
Study strength low
not reported
0.06
The qualitative finding that reinforcement-learning firms can produce tacitly collusive wage outcomes is robust across variations in platform design, worker heterogeneity, and worker search behavior. Market Structure negative Persistence of collusive wage-setting outcomes across model specifications
Reading fidelity high
Study strength low
not reported
0.06
The simulated collusive wage-setting behavior harms workers by producing wage outcomes different from the competitive benchmark. Wages negative Posted wages and worker welfare relative to the competitive benchmark
Reading fidelity medium
Study strength low
not reported
0.04

Notes