0 cumulative citations
View corpus contextAn XGBoost model trained on nine U.S. national workforce datasets predicts employee turnover with 92.8% accuracy, and SHAP analysis flags tenure, pay, age, benefits and local job openings as the dominant predictors; however, the study is predictive rather than causal and leaves key harmonization and validation details unreported.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextEmployee turnover remains one of the most significant workforce challenges affecting organizational productivity, competitiveness, and long-term sustainability. This study examined national employee turnover trends in the United States between 2010 and 2025 and comparatively evaluated the performance of machine learning algorithms for predicting employee turnover using harmonized data from nine nationally representative workforce datasets: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, and SIPP. A comparative quantitative design integrating descriptive trend analysis, supervised machine learning, and Explainable Artificial Intelligence (SHAP) was employed. Seven classification algorithms: Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, Artificial Neural Network, Gradient Boosting Machine, and Extreme Gradient Boosting (XGBoost), were evaluated using cross-validation and multiple performance metrics. The results revealed three distinct labor market phases: post-recession recovery (2010–2019), COVID-19 disruption (2020), and post-pandemic adjustment (2021–2025), with workforce mobility remaining above historical levels despite recent stabilization. XGBoost achieved the highest predictive accuracy (92.8%), outperforming all other models, followed by Random Forest (91.6%) and Gradient Boosting (90.9%). SHAP analysis identified organizational tenure, annual wage, age, employee benefits, and regional job openings as the most influential predictors of employee turnover, while employment history and compensation emerged as the dominant predictor domains. The findings demonstrate that employee turnover is driven by the interaction of organizational, demographic, occupational, and labor market factors rather than isolated organizational characteristics. The study concludes that integrating nationally representative workforce data with machine learning and explainable artificial intelligence provides a robust framework for proactive employee retention, strategic workforce planning, and evidence-based labor market policy.
Summary
Main Finding
Using harmonized national U.S. workforce data (2010–2025) and explainable machine learning, the study shows that employee turnover is best predicted by ensemble tree methods (XGBoost achieved 92.8% accuracy), and that turnover is driven by interacting organizational, demographic, occupational, compensation, and labor‑market factors rather than single internal HR characteristics. SHAP analysis identified organizational tenure, annual wage, age, employee benefits, and regional job openings as the most influential predictors; at the domain level, employment history and compensation dominated.
Key Points
-
Data span and focus
- Nationally representative integration of nine U.S. workforce datasets covering 2010–2025: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, SIPP.
- Analysis emphasizes broad labor‑market context: three phases identified — post‑recession recovery (2010–2019), COVID‑19 disruption (2020), post‑pandemic adjustment (2021–2025) with workforce mobility remaining above many pre‑pandemic norms.
-
Methods and modelling
- Comparative supervised classification of turnover using 7 algorithms: Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, Artificial Neural Network, Gradient Boosting Machine, XGBoost.
- Standardized preprocessing, hyperparameter optimization, and cross‑validation with multiple performance metrics for fair comparison.
- Explainable AI: SHAP (SHapley Additive exPlanations) used to quantify variable importance and to decompose contributions of predictor domains.
-
Performance and interpretability
- Top performers: XGBoost (92.8% accuracy), Random Forest (91.6%), Gradient Boosting (90.9%).
- SHAP identifies top individual predictors: organizational tenure, annual wage, age, employee benefits, regional job openings.
- Predictor domains ranked: employment history and compensation & benefits most influential; demographic, occupational, and labor‑market features also important via interactions.
-
Contributions
- Demonstrates feasibility and value of integrating multiple national surveys for generalizable turnover prediction.
- Combines trend analysis, high‑accuracy prediction, and XAI to inform both firm-level retention strategies and public workforce policy.
Data & Methods
- Data sources: JOLTS, CPS, ACS, NLSY97, LEHD, NCS, OEWS, O*NET, SIPP — harmonized to produce a unified feature set capturing demographics, employment history, wages/benefits, occupational attributes, and regional labor‑market measures.
- Outcome: employee turnover (voluntary/involuntary separations as defined from survey harmonization).
- Modeling pipeline:
- Feature engineering across domains (tenure, compensation, benefits, skill/occupation indices, regional job openings, prior employment history, demographics).
- Standard preprocessing (imputation, encoding, scaling where appropriate).
- Hyperparameter tuning and k‑fold cross‑validation; evaluation with multiple metrics (accuracy reported as headline metric; other metrics used though not detailed in the abstract).
- Explainability: global and local SHAP analyses to rank predictors and interpret model decisions.
- Robustness considerations noted by authors:
- Use of nationally representative, longitudinal sources to improve external validity.
- Consistent evaluation protocol across algorithms to reduce methodological bias in comparisons.
Implications for AI Economics
-
For firms and HR analytics
- High predictive performance of tree‑ensemble models (XGBoost/RF/GBM) implies practical tools for proactive retention — targeted interventions, optimized training investments, and more efficient allocation of retention spending.
- Explainability (SHAP) makes models actionable and defensible to stakeholders, helping translate predictions into policy (e.g., tenure‑based interventions, benefit redesign).
-
For labor markets and policy
- Predictive workforce tools based on national data can inform regional workforce planning, identify sectors with elevated churn risk, and guide public investments (training, unemployment supports, mobility programs).
- Widespread adoption of such tools could alter firm behavior and labor bargaining: better retention forecasting may reduce recruitment costs but could also enable stronger employer strategies that affect wage/offer dynamics and, in aggregate, influence labor supply and mobility.
-
For AI economics research and deployment
- Demonstrates value of combining multi‑source, representative datasets with XAI to produce generalizable, interpretable predictions — a model for future macro‑level AI economic applications.
- Raises important caveats that affect economic interpretation and policy:
- Prediction vs causation: high predictive importance (SHAP) does not establish causal effects; policies should be informed by causal analyses or experimentation before large interventions.
- Temporal nonstationarity: labor markets evolve (e.g., pandemic shocks); models must be continuously updated and monitored for performance drift.
- Equity and fairness: using demographic and employment data to predict turnover risks can propagate or amplify bias (e.g., differential monitoring or interventions by group). Fairness constraints and regulatory compliance are necessary.
- Privacy and governance: integrating and operationalizing national data for firm use presents data‑protection and governance challenges.
- Market concentration risks: if only some firms adopt powerful predictive retention tools, competitive dynamics and labor market power could shift (monopsony concerns).
-
Research directions
- Combine predictive models with causal inference (experiments, quasi‑experimental designs) to identify effective retention policies.
- Evaluate distributional effects of predictive HR tools across worker groups and regions.
- Study equilibrium effects: how firm adoption of turnover prediction alters wages, vacancy dynamics, and aggregate mobility.
Limitations noted or implied - Harmonization of multiple surveys improves external validity but entails measurement and harmonization error. - Label definitions and survey timing may limit precision (e.g., distinguishing voluntary vs involuntary separations). - The study is predictive; causal interpretation requires further work. - Ethical/regulatory risks require mitigation before deployment in personnel decisions.
Summary conclusion - Integrating national workforce data with ensemble machine learning and XAI yields high predictive accuracy and interpretable drivers of turnover. This approach promises efficiency gains for firms and actionable insights for policymakers, but careful attention to causality, fairness, temporal robustness, and governance is required for responsible economic and labor‑market applications.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| U.S. employee turnover from 2010 to 2025 followed three labor-market phases: post-recession recovery (2010–2019), COVID-19 disruption (2020), and post-pandemic adjustment (2021–2025). Turnover | mixed | Employee turnover and workforce mobility trends over time |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Workforce mobility remained above historical levels during the post-pandemic period despite recent stabilization. Turnover | positive | Workforce mobility and employee turnover levels |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Among the seven evaluated machine-learning algorithms, XGBoost achieved the highest employee-turnover predictive accuracy at 92.8%. Turnover | positive | Predictive accuracy for employee turnover classification |
Reading fidelity
high
Study strength
medium
|
92.8% predictive accuracy
|
| Random Forest achieved 91.6% predictive accuracy and Gradient Boosting achieved 90.9% predictive accuracy in employee-turnover classification, ranking below XGBoost. Turnover | positive | Predictive accuracy for employee turnover classification |
Reading fidelity
high
Study strength
medium
|
Random Forest: 91.6% accuracy; Gradient Boosting: 90.9% accuracy
|
| Organizational tenure, annual wage, age, employee benefits, and regional job openings were the most influential individual predictors of employee turnover according to SHAP analysis. Turnover | mixed | Relative predictor importance for employee turnover |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Employment history and compensation were the dominant predictor domains for employee turnover. Turnover | mixed | Relative importance of predictor domains for employee turnover |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Employee turnover is driven by interactions among organizational, demographic, occupational, and labor-market factors rather than by isolated organizational characteristics. Turnover | mixed | Determinants of employee turnover |
Reading fidelity
high
Study strength
medium
|
not reported
|