The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

An XGBoost model trained on nine U.S. national workforce datasets predicts employee turnover with 92.8% accuracy, and SHAP analysis flags tenure, pay, age, benefits and local job openings as the dominant predictors; however, the study is predictive rather than causal and leaves key harmonization and validation details unreported.

PREDICTING EMPLOYEE TURNOVER IN THE UNITED STATES: A COMPARATIVE MACHINE LEARNING ANALYSIS USING NATIONAL WORKFORCE SURVEY DATA (2010–2025)
Tarisai Blessing Makaba, Charity Varaidzo Katokwe, Gamuchirai Thomas Hlatywayo, Tanatswa Esther Nyoni · August 20, 2026 · Magna Scientia Advanced Research and Reviews
openalex descriptive low evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Tarisai Blessing Makaba provider ID
  2. Charity Varaidzo Katokwe provider ID
  3. Gamuchirai Thomas Hlatywayo provider ID
  4. Tanatswa Esther Nyoni provider ID

Semantic Scholar

Latest observation:

  1. Tarisai Blessing Makaba provider ID
  2. Charity Varaidzo Katokwe provider ID
  3. Gamuchirai Thomas Hlatywayo provider ID
  4. Tanatswa Esther Nyoni provider ID
Using harmonized U.S. national workforce data from 2010–2025, the authors find that XGBoost predicts employee turnover best (92.8% accuracy) and that tenure, wages, age, benefits and regional job openings are the most important predictors according to SHAP.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Employee turnover remains one of the most significant workforce challenges affecting organizational productivity, competitiveness, and long-term sustainability. This study examined national employee turnover trends in the United States between 2010 and 2025 and comparatively evaluated the performance of machine learning algorithms for predicting employee turnover using harmonized data from nine nationally representative workforce datasets: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, and SIPP. A comparative quantitative design integrating descriptive trend analysis, supervised machine learning, and Explainable Artificial Intelligence (SHAP) was employed. Seven classification algorithms: Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, Artificial Neural Network, Gradient Boosting Machine, and Extreme Gradient Boosting (XGBoost), were evaluated using cross-validation and multiple performance metrics. The results revealed three distinct labor market phases: post-recession recovery (2010–2019), COVID-19 disruption (2020), and post-pandemic adjustment (2021–2025), with workforce mobility remaining above historical levels despite recent stabilization. XGBoost achieved the highest predictive accuracy (92.8%), outperforming all other models, followed by Random Forest (91.6%) and Gradient Boosting (90.9%). SHAP analysis identified organizational tenure, annual wage, age, employee benefits, and regional job openings as the most influential predictors of employee turnover, while employment history and compensation emerged as the dominant predictor domains. The findings demonstrate that employee turnover is driven by the interaction of organizational, demographic, occupational, and labor market factors rather than isolated organizational characteristics. The study concludes that integrating nationally representative workforce data with machine learning and explainable artificial intelligence provides a robust framework for proactive employee retention, strategic workforce planning, and evidence-based labor market policy.

Summary

Main Finding

Using harmonized national U.S. workforce data (2010–2025) and explainable machine learning, the study shows that employee turnover is best predicted by ensemble tree methods (XGBoost achieved 92.8% accuracy), and that turnover is driven by interacting organizational, demographic, occupational, compensation, and labor‑market factors rather than single internal HR characteristics. SHAP analysis identified organizational tenure, annual wage, age, employee benefits, and regional job openings as the most influential predictors; at the domain level, employment history and compensation dominated.

Key Points

  • Data span and focus

    • Nationally representative integration of nine U.S. workforce datasets covering 2010–2025: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, SIPP.
    • Analysis emphasizes broad labor‑market context: three phases identified — post‑recession recovery (2010–2019), COVID‑19 disruption (2020), post‑pandemic adjustment (2021–2025) with workforce mobility remaining above many pre‑pandemic norms.
  • Methods and modelling

    • Comparative supervised classification of turnover using 7 algorithms: Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, Artificial Neural Network, Gradient Boosting Machine, XGBoost.
    • Standardized preprocessing, hyperparameter optimization, and cross‑validation with multiple performance metrics for fair comparison.
    • Explainable AI: SHAP (SHapley Additive exPlanations) used to quantify variable importance and to decompose contributions of predictor domains.
  • Performance and interpretability

    • Top performers: XGBoost (92.8% accuracy), Random Forest (91.6%), Gradient Boosting (90.9%).
    • SHAP identifies top individual predictors: organizational tenure, annual wage, age, employee benefits, regional job openings.
    • Predictor domains ranked: employment history and compensation & benefits most influential; demographic, occupational, and labor‑market features also important via interactions.
  • Contributions

    • Demonstrates feasibility and value of integrating multiple national surveys for generalizable turnover prediction.
    • Combines trend analysis, high‑accuracy prediction, and XAI to inform both firm-level retention strategies and public workforce policy.

Data & Methods

  • Data sources: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, SIPP — harmonized to produce a unified feature set capturing demographics, employment history, wages/benefits, occupational attributes, and regional labor‑market measures.
  • Outcome: employee turnover (voluntary/involuntary separations as defined from survey harmonization).
  • Modeling pipeline:
    • Feature engineering across domains (tenure, compensation, benefits, skill/occupation indices, regional job openings, prior employment history, demographics).
    • Standard preprocessing (imputation, encoding, scaling where appropriate).
    • Hyperparameter tuning and k‑fold cross‑validation; evaluation with multiple metrics (accuracy reported as headline metric; other metrics used though not detailed in the abstract).
    • Explainability: global and local SHAP analyses to rank predictors and interpret model decisions.
  • Robustness considerations noted by authors:
    • Use of nationally representative, longitudinal sources to improve external validity.
    • Consistent evaluation protocol across algorithms to reduce methodological bias in comparisons.

Implications for AI Economics

  • For firms and HR analytics

    • High predictive performance of tree‑ensemble models (XGBoost/RF/GBM) implies practical tools for proactive retention — targeted interventions, optimized training investments, and more efficient allocation of retention spending.
    • Explainability (SHAP) makes models actionable and defensible to stakeholders, helping translate predictions into policy (e.g., tenure‑based interventions, benefit redesign).
  • For labor markets and policy

    • Predictive workforce tools based on national data can inform regional workforce planning, identify sectors with elevated churn risk, and guide public investments (training, unemployment supports, mobility programs).
    • Widespread adoption of such tools could alter firm behavior and labor bargaining: better retention forecasting may reduce recruitment costs but could also enable stronger employer strategies that affect wage/offer dynamics and, in aggregate, influence labor supply and mobility.
  • For AI economics research and deployment

    • Demonstrates value of combining multi‑source, representative datasets with XAI to produce generalizable, interpretable predictions — a model for future macro‑level AI economic applications.
    • Raises important caveats that affect economic interpretation and policy:
      • Prediction vs causation: high predictive importance (SHAP) does not establish causal effects; policies should be informed by causal analyses or experimentation before large interventions.
      • Temporal nonstationarity: labor markets evolve (e.g., pandemic shocks); models must be continuously updated and monitored for performance drift.
      • Equity and fairness: using demographic and employment data to predict turnover risks can propagate or amplify bias (e.g., differential monitoring or interventions by group). Fairness constraints and regulatory compliance are necessary.
      • Privacy and governance: integrating and operationalizing national data for firm use presents data‑protection and governance challenges.
      • Market concentration risks: if only some firms adopt powerful predictive retention tools, competitive dynamics and labor market power could shift (monopsony concerns).
  • Research directions

    • Combine predictive models with causal inference (experiments, quasi‑experimental designs) to identify effective retention policies.
    • Evaluate distributional effects of predictive HR tools across worker groups and regions.
    • Study equilibrium effects: how firm adoption of turnover prediction alters wages, vacancy dynamics, and aggregate mobility.

Limitations noted or implied - Harmonization of multiple surveys improves external validity but entails measurement and harmonization error. - Label definitions and survey timing may limit precision (e.g., distinguishing voluntary vs involuntary separations). - The study is predictive; causal interpretation requires further work. - Ethical/regulatory risks require mitigation before deployment in personnel decisions.

Summary conclusion - Integrating national workforce data with ensemble machine learning and XAI yields high predictive accuracy and interpretable drivers of turnover. This approach promises efficiency gains for firms and actionable insights for policymakers, but careful attention to causality, fairness, temporal robustness, and governance is required for responsible economic and labor‑market applications.

Assessment

Paper Typedescriptive Evidence Strengthlow — The paper reports strong out-of-sample predictive performance (XGBoost 92.8%) and uses multiple national datasets, but it provides no credible causal identification strategy—results are predictive/correlational only. The abstract and introduction do not report key validation details (temporal holdout, external validation, sample sizes after harmonization, treatment of missing data, class balance), so predictive results may be optimistic or not robust to deployment settings. Methods Rigormedium — The study uses an appropriate comparative ML framework (multiple algorithms, cross-validation, hyperparameter tuning, and SHAP for interpretability) and integrates many nationally representative data sources, which strengthens external relevance. However important methodological details are missing or unclear in the supplied text: how datasets were harmonized and linked, the precise outcome definition (voluntary vs all separations), sample construction and size, treatment of missingness, class imbalance, time-based validation to avoid leakage, and robustness checks, all of which are crucial for trustworthy predictive modelling. SampleHarmonized records drawn from nine U.S. national workforce datasets (JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, SIPP) covering 2010–2025; exact merged sample size, unit of observation (person-month, person-year, or aggregated cells), linkage procedure, and post-harmonization sample composition are not reported in the supplied text. Themeslabor_markets human_ai_collab GeneralizabilityGeographic: limited to the United States and its institutional/labor market context, Temporal: trained on 2010–2025 data; may not generalize to future structural changes or shocks after 2025, Measurement/harmonization: merging heterogeneous surveys risks inconsistent variable definitions and measurement error, Outcome definition: unclear whether turnover refers to voluntary quits only or includes involuntary separations, limiting interpretability, Deployment: high reported accuracy may not transfer to firm-level operational settings without local recalibration, Sample selection: survey nonresponse, panel attrition (e.g., NLSY, SIPP) and linkage restrictions may bias estimates

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
U.S. employee turnover from 2010 to 2025 followed three labor-market phases: post-recession recovery (2010–2019), COVID-19 disruption (2020), and post-pandemic adjustment (2021–2025). Turnover mixed Employee turnover and workforce mobility trends over time
Reading fidelity high
Study strength medium
not reported
0.18
Workforce mobility remained above historical levels during the post-pandemic period despite recent stabilization. Turnover positive Workforce mobility and employee turnover levels
Reading fidelity high
Study strength medium
not reported
0.18
Among the seven evaluated machine-learning algorithms, XGBoost achieved the highest employee-turnover predictive accuracy at 92.8%. Turnover positive Predictive accuracy for employee turnover classification
Reading fidelity high
Study strength medium
92.8% predictive accuracy
0.18
Random Forest achieved 91.6% predictive accuracy and Gradient Boosting achieved 90.9% predictive accuracy in employee-turnover classification, ranking below XGBoost. Turnover positive Predictive accuracy for employee turnover classification
Reading fidelity high
Study strength medium
Random Forest: 91.6% accuracy; Gradient Boosting: 90.9% accuracy
0.18
Organizational tenure, annual wage, age, employee benefits, and regional job openings were the most influential individual predictors of employee turnover according to SHAP analysis. Turnover mixed Relative predictor importance for employee turnover
Reading fidelity high
Study strength medium
not reported
0.18
Employment history and compensation were the dominant predictor domains for employee turnover. Turnover mixed Relative importance of predictor domains for employee turnover
Reading fidelity high
Study strength medium
not reported
0.18
Employee turnover is driven by interactions among organizational, demographic, occupational, and labor-market factors rather than by isolated organizational characteristics. Turnover mixed Determinants of employee turnover
Reading fidelity high
Study strength medium
not reported
0.18

Notes