The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Using ML decision aids can erode workers' decision-making skills: randomized experiments find sizable performance drops when the system is removed, particularly among users who put greater trust in the tool.

The Dependency Dilemma: How Machine Learning Decision Aids can Undermine Skill Growth
Kevin Bauer, Michael Nofer, Benjamin Henrich, Hendrik Drachsler, Oliver Hinz · July 20, 2026 · Business & Information Systems Engineering
openalex rct high evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Kevin Bauer provider ID
  2. Michael Nofer provider ID
  3. Benjamin Henrich provider ID
  4. Hendrik Drachsler provider ID
  5. Oliver Hinz provider ID

Semantic Scholar

Latest observation:

  1. Kevin Bauer provider ID
  2. Michael Nofer provider ID
  3. Benjamin Henrich provider ID
  4. Hendrik Drachsler provider ID
  5. Oliver Hinz provider ID
Randomized evidence shows that relying on ML decision aids impairs the development of human decision-making skills, producing significant performance declines when the system is unavailable, with greater trust amplifying the deficit.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract Advances in Machine Learning (ML) have led organizations to increasingly implement ML decision aids to enhance employees’ decision-making performance. While such systems can improve organizational efficiency in many contexts, they may inadvertently impact the development of human decision-making skills. Drawing on cognitive theories, this study examines how the use of ML decision aids impact skill development and performance. Using a novel experimental design tailored to address organizational challenges and endogeneity concerns, the study identifies causal effects of reliance on ML predictions on skill development in decision making. Specifically, it is demonstrated that reliance on ML predictions in a prediction-making task can hinder the development of critical decision-making skills, resulting in significant performance drops when the system becomes unavailable. Furthermore, it is found that the extent of trust in the system's predictions strongly influences the severity of this skill deficit. These findings highlight the need for thoughtful integration of ML decision aids, emphasizing the importance of balancing reliance with skill retention to mitigate risks associated with temporary or permanent system disruptions.

Summary

Main Finding

Reliance on ML decision aids can causally undermine the development of human decision-making skills. In an incentivized experiment, participants who used ML predictions learned less and showed significant performance drops when the aid was removed; the greater the trust/reliance on the ML outputs, the larger the later skill deficit.

Key Points

  • ML decision aids improve short-term decision performance but can produce "deskilling": users stop actively engaging with task-relevant information and fail to form robust procedural knowledge.
  • The paper frames learning via the ACT‑R cognitive model: skills develop by converting declarative knowledge (observations, feedback) into procedural production rules through practice. Excessive reliance on ML interrupts this conversion.
  • Skill deficits become visible when the system is unavailable (intentional discontinuance, malfunction, or rapid concept drift), since users lack the procedural rules to perform unaided.
  • Trust in the ML system strongly moderates the effect: higher trust → more reliance → larger decline in unaided performance.
  • Concept drift and sudden, hard-to-anticipate environment shifts (e.g., pandemics) increase the value of retained human skills because ML models may rapidly lose validity; even advanced MLOps cannot always prevent sudden failures.
  • The study builds on deskilling/automation literature but focuses specifically on ML decision aids (as distinct from older expert systems) and the acquisition of entirely new cognitive/problem-solving skills.

Data & Methods

  • Primary method: a novel, incentivized laboratory experiment designed to identify causal effects and to mitigate endogeneity concerns.
  • Task: participants solved logical puzzles where discovering the underlying rule/strategy (procedural skill) increased earnings. Continuous feedback enabled trial-and-error learning.
  • Treatment vs control: some participants received ML-generated predictions (decision aid) while others did not. The aid’s error rates were observable to participants (allowing analysis of trust formation).
  • Identification strategy: causal inference comes from the experimental assignment and the deliberate removal of the ML aid (system discontinuance) to observe downstream effects on unaided performance.
  • Measured outcomes: reliance/trust in the ML predictions, learning/skill acquisition during training, and performance after the ML aid was removed (performance drop as indicator of deskilling).
  • Note: exact sample size, demographic composition, ML model type, and statistical estimates are not included in the provided excerpt.

Implications for AI Economics

  • Human capital depreciation under automation: Firms adopting ML aids face a trade-off between immediate productivity gains and longer‑term erosion of employee skills. Economic models of automation should account for endogenous human-skill depreciation and its effect on organizational resilience.
  • Resilience and risk valuation: The value of human skills increases with the risk of system failure or sudden concept drift. Firms operating in high‑drift environments (fast-changing markets, crisis-prone sectors) should discount returns from automation more heavily or invest more in maintaining skills.
  • Designing incentives and institutions: To mitigate deskilling, organizations should (a) design ML aids that encourage active human engagement (e.g., require justification, promote counterfactual thinking), (b) rotate tasks or periodically remove aids to force practice, (c) maintain training programs and human-in-the-loop interventions, and (d) monitor trust formation and reliance to detect unhealthy dependence.
  • Product design and regulation: ML system designers and regulators should consider requirements that preserve skill retention (transparency, adjustable automation levels, auditability). Policies encouraging human oversight and retraining subsidies can reduce societal costs of deskilling.
  • Labor markets and welfare: Short-term employment productivity gains may mask longer-term reductions in worker versatility and employability. Policymakers should factor deskilling externalities into assessments of AI adoption and labor-market interventions.
  • Cost–benefit and investment decisions: Firms must internalize the expected cost of potential discontinuance (retraining, temporary productivity loss). Dynamic investment decisions (how much to automate vs. keep manual) should incorporate the probability and cost of skill erosion and future model failures.
  • Research agenda for AI economics: Empirical and theoretical models should quantify skill depreciation rates under different automation regimes, characterize optimal hybrid policies (when to automate vs. preserve practice), and estimate macroeconomic effects of widespread deskilling (e.g., on adaptability, innovation).

If you want, I can (a) draft a short formal model of the trade‑off between immediate gains from ML aids and long-run skill depreciation, or (b) extract specific managerial recommendations for ML deployment (e.g., UI features, training schedules) based on the paper.

Assessment

Paper Typerct Evidence Strengthhigh — Causal identification is achieved via random assignment and a within-subjects removal test that isolates skill acquisition from on-the-job performance while the aid is present; this directly links reliance on ML to later performance drops, supporting a causal interpretation. Methods Rigorhigh — The study uses a preplanned experimental intervention addressing endogeneity, objective performance measures, and heterogeneity analysis by trust; however, the abstract does not report sample size, field vs. lab setting, or robustness checks, which tempers assessment of overall rigor. SampleHuman subjects performing a prediction-making task in an experimental setting; participants were randomly assigned to use ML predictions or not, with performance recorded during aid availability and after removal; the abstract does not specify whether subjects were students, crowdworkers, or employees, nor sample size or country. Themeshuman_ai_collab skills_training productivity org_design IdentificationRandomized experimental design in which participants were assigned to use an ML decision aid (vs. control/no-aid) during a prediction-making task; performance was measured both while the aid was available and after it was removed to isolate effects on skill development, with trust in the system measured (and plausibly manipulated or instrumented) to explain heterogeneity. GeneralizabilityLab/experimental task may not reflect complexity of real-world organizational decisions, Participant pool unspecified (e.g., students or MTurk) may limit applicability to professional workers, Short-term experimental horizon may not capture long-run learning or on-the-job training dynamics, Findings may depend on task type and ML system accuracy/sophistication and thus not generalize across domains or firm contexts, Cultural, institutional, or industry differences in reliance/trust not addressed

Claims (4)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Reliance on ML predictions in a prediction-making task can hinder the development of critical decision-making skills. Skill Acquisition negative development of decision-making skills
Reading fidelity high
Study strength medium
not reported
0.6
Reliance on ML predictions results in significant performance drops when the system becomes unavailable. Decision Quality negative performance on the prediction-making task when ML system unavailable
Reading fidelity high
Study strength medium
not reported
0.6
The extent of trust in the system's predictions strongly influences the severity of this skill deficit. Skill Obsolescence negative severity of skill deficit as a function of trust in system predictions
Reading fidelity high
Study strength medium
not reported
0.6
These findings highlight the need for thoughtful integration of ML decision aids, emphasizing the importance of balancing reliance with skill retention to mitigate risks associated with temporary or permanent system disruptions. Organizational Efficiency positive organizational risk mitigation / skill retention practices
Reading fidelity medium
Study strength speculative
not reported
0.06

Notes