0 cumulative citations
View corpus contextReinforcement‑learning employers on online labor platforms can tacitly collude on wages — but that risk is concentrated when few firms compete and disappears if common training techniques like experience replay are used.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextOnline digital labor platforms are increasingly important as tools for matching employers and workers. They reduce labor market frictions and provide the opportunity for remote work, but concerns have been raised about whether workers receive adequate incomes. We study how the use of machine-learning based algorithms for the determination of wage offers affects workers’ wages on online labor platforms by setting up a simulation framework and analyzing outcomes under different assumptions about platform and algorithm design. Firms use reinforcement-learning to update posted wages on the platform, and heterogeneous workers send applications based on the posted information. We show that, if firms use a deep Q-network (DQN), a state-of-the-art machine learning algorithm, collusive wages emerge when the number of competitors is small, and no experience replay is used. Our findings are robust across many features of the model, including the design of the online labor platform.
Summary
Main Finding
When firms on online labor platforms use state-of-the-art reinforcement learning (a deep Q-network, DQN) to set posted wages, they can learn tacitly collusive wage-setting behavior that harms workers — but this outcome is conditional. Collusive (non-competitive) wage outcomes emerge robustly in simulation when the number of competing firms is small and when a common DQN training technique (experience replay) is not used. The qualitative result holds across many variants of platform and model design.
Key Points
- Setting: repeated matching on online digital labor platforms where firms post wages and heterogeneous workers apply based on posted information.
- Strategic learning: firms independently update posted wages using reinforcement learning (DQN); there is no explicit coordination or communication.
- Main mechanism: with small numbers of competitors and without experience replay, the learning dynamics converge to tacitly collusive wage policies (i.e., joint outcomes different from the competitive benchmark).
- Role of experience replay: including experience replay in DQN training disrupts the temporal correlations that facilitate collusive dynamics and prevents the emergence of collusion in the authors’ simulations.
- Competition intensity matters: as the number of competitors increases, collusive outcomes are less likely; the effect depends on market structure.
- Robustness: the collusion result is robust to many modeled features, including different platform designs and worker heterogeneity/search behavior.
Data & Methods
- Approach: agent-based simulation / multi-agent reinforcement learning framework (no empirical field data).
- Agents:
- Firms: use reinforcement learning (deep Q-network, DQN) to choose posted wage offers over repeated rounds.
- Workers: heterogeneous in characteristics and application behavior; they respond to posted wages and other posted information in the matching process.
- Experiments:
- Vary number of competing firms.
- Toggle DQN features (notably experience replay vs. no replay).
- Vary platform design parameters and worker-side primitives (heterogeneity, search frictions).
- Measure outcomes such as equilibrium wages, firm profits, worker application patterns, and convergence properties.
- Diagnostics: analyze learned policies and dynamics to identify when tacit collusion emerges and under what algorithmic/design conditions it fails to appear.
Implications for AI Economics
- Algorithmic tacit collusion extends beyond product pricing: employer-side use of RL can generate anti-competitive wage outcomes in labor markets.
- Algorithm design choices matter for market outcomes: low-level ML implementation details (e.g., experience replay) can be policy-relevant levers that either facilitate or prevent collusive dynamics.
- Regulatory and platform policy:
- Antitrust and labor regulators should consider algorithmic learning dynamics as potential sources of coordinated (tacit) employer behavior.
- Platforms could mitigate risk by encouraging or requiring algorithmic features that break temporal correlations (e.g., experience replay, enforced randomization) or by auditing posted-wage dynamics.
- Increasing market competitiveness (more employers) and transparency of posting practices may reduce the likelihood of collusion.
- Research directions:
- Empirical validation on real platforms to detect signatures of algorithmic collusion in posted wages.
- Study of other algorithmic architectures, information structures, and multi-agent learning protocols.
- Designing regulation-friendly algorithmic designs and platform rules that preserve market efficiency and worker welfare.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In agent-based simulations of online labor platforms, firms using deep Q-networks to set posted wages can learn tacitly collusive wage-setting behavior without explicit coordination or communication. Market Structure | negative | Whether firms converge to non-competitive, tacitly collusive wage policies |
Reading fidelity
high
Study strength
low
|
not reported
|
| Collusive wage outcomes emerge robustly in the simulations when the number of competing firms is small and experience replay is omitted from DQN training. Market Structure | negative | Emergence of collusive or non-competitive posted-wage outcomes |
Reading fidelity
high
Study strength
low
|
not reported
|
| Including experience replay in DQN training disrupts the temporal correlations associated with collusive dynamics and prevents collusion from emerging in the authors' simulations. Market Structure | positive | Emergence or prevention of tacit wage collusion |
Reading fidelity
high
Study strength
low
|
not reported
|
| Increasing the number of competing firms makes collusive wage outcomes less likely. Market Structure | positive | Likelihood of tacitly collusive wage-setting behavior |
Reading fidelity
high
Study strength
low
|
not reported
|
| The qualitative finding that reinforcement-learning firms can produce tacitly collusive wage outcomes is robust across variations in platform design, worker heterogeneity, and worker search behavior. Market Structure | negative | Persistence of collusive wage-setting outcomes across model specifications |
Reading fidelity
high
Study strength
low
|
not reported
|
| The simulated collusive wage-setting behavior harms workers by producing wage outcomes different from the competitive benchmark. Wages | negative | Posted wages and worker welfare relative to the competitive benchmark |
Reading fidelity
medium
Study strength
low
|
not reported
|