The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Probabilistic boosting outperforms deep tabular nets for stable credit-risk probabilities: NGBoost–SHAP gives better-calibrated, less regime-sensitive 1‑year bankruptcy forecasts across US and China, whereas TabNet–SHAP adapts faster but is more variable and preprocessing-sensitive.

Predicting Financial Distress: A Comparative Analysis of Explainable Machine Learning Models for Cross‐Market Bankruptcy Prediction
Aqsa Bilal Hussain, Jingchun Sun · August 31, 2026 · Journal of Forecasting
openalex correlational medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Aqsa Bilal Hussain provider ID
  2. Jingchun Sun provider ID

Semantic Scholar

Latest observation:

  1. Aqsa Bilal Hussain unresolved corpus identity
  2. Jing-Chun Sun unresolved corpus identity
Across U.S. and Chinese listed firms (2015–2022) NGBoost with SHAP yields more stable and better-calibrated 1‑year bankruptcy probabilities across the COVID regime break, while TabNet with SHAP is more adaptive but produces more variable, regime‑sensitive results.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

ABSTRACT Despite substantial advances in machine‐learning approaches to bankruptcy prediction, less is known about whether explainable models remain reliable across markets and during structural breaks. This study examines that issue by comparing two explainable architectures, TabNet–SHAP and NGBoost–SHAP, for 1‐year‐ahead bankruptcy prediction in the United States and China during 2015–2022. Using a balanced panel of listed firms from both markets, we evaluate out‐of‐sample performance before and after COVID‐19 and benchmark both models against conventional linear and ensemble classifiers. Model uncertainty is assessed through cross‐validation and bootstrap confidence intervals. The results show that NGBoost–SHAP provides more stable and better‐calibrated probabilistic predictions across regimes, whereas TabNet–SHAP produces more variable results and is more sensitive to regime shifts and preprocessing choices. We interpret this pattern as an architecture‐specific trade‐off between stability and adaptability rather than as an unconditional ranking of the two models. SHAP‐based interpretation indicates that profitability and capital‐structure variables account for most of the explanatory signal within each market and model. The post‐COVID Chinese subsample shows the most concentrated feature‐attribution structure, although SHAP magnitudes are interpreted only as within‐model rankings rather than as directly comparable cross‐market quantities. The study contributes by proposing an adaptability–stability–interpretability framework for selecting bankruptcy prediction models under different institutional and market conditions. It also provides evidence that explainable ensemble and deep tabular models can recover economically meaningful distress signals across markets while differing in their calibration, stability, and sensitivity to economic regime change.

Summary

Main Finding

NGBoost–SHAP yields more stable and better‑calibrated 1‑year bankruptcy probabilities across the U.S. and Chinese markets and across the COVID regime break, whereas TabNet–SHAP is more adaptive but produces more variable, regime‑sensitive results. This pattern reflects an architecture‑specific trade‑off between stability and adaptability rather than an absolute ranking of the two models.

Key Points

  • Models compared: TabNet–SHAP (deep tabular architecture with SHAP explanations) and NGBoost–SHAP (probabilistic boosting with SHAP).
  • Scope: 1‑year‑ahead bankruptcy prediction for listed firms in the United States and China, 2015–2022, using a balanced panel.
  • Evaluation: out‑of‑sample performance before and after the COVID shock, benchmarked against conventional linear and ensemble classifiers.
  • Uncertainty quantification: cross‑validation and bootstrap confidence intervals used to assess model uncertainty and stability.
  • Main empirical results:
    • NGBoost–SHAP produced more stable predictions across regimes and better probabilistic calibration.
    • TabNet–SHAP showed greater variability, higher sensitivity to regime shifts and preprocessing choices, and less stable calibration.
  • Explainability findings (SHAP):
    • Profitability and capital‑structure variables carry most of the explanatory signal in both markets and models.
    • The post‑COVID Chinese subsample exhibited the most concentrated feature‑attribution pattern.
    • SHAP values are treated as within‑model feature rankings and are not directly comparable across models or markets.

Data & Methods

  • Data: Balanced panel of publicly listed firms in the U.S. and China covering 2015–2022.
  • Prediction target: 1‑year ahead bankruptcy/distress.
  • Models evaluated:
    • TabNet with SHAP for local/global feature attributions.
    • NGBoost with SHAP to provide probabilistic forecasts and feature attributions.
    • Conventional linear and ensemble baselines for benchmarking.
  • Validation strategy:
    • Out‑of‑sample tests split around the COVID regime break (pre/post).
    • Cross‑validation for model selection and stability checks.
    • Bootstrap procedures to compute confidence intervals for performance metrics and SHAP summaries.
  • Interpretability: SHAP used to rank feature importance within each model; magnitudes interpreted cautiously as relative (within‑model) signals.

Implications for AI Economics

  • Model selection under regime change:
    • Use NGBoost‑like probabilistic ensembles when stability and well‑calibrated risk probabilities are critical (e.g., regulatory stress tests, credit scoring under varying macro states).
    • Consider TabNet‑like deep tabular models for environments where adaptability to new patterns may pay off, but pair them with rigorous validation and robustness checks.
  • Model monitoring and governance:
    • Regular out‑of‑sample testing across economic regimes and bootstrap uncertainty quantification are essential to detect calibration drift and preprocessing sensitivity.
    • SHAP summaries are useful to recover economically meaningful drivers (profitability, leverage) but should not be used to compare importance magnitudes across models or markets without caution.
  • Policy and risk management:
    • Explainable ML can recover relevant distress signals across institutional contexts, supporting cross‑market credit surveillance, but differing calibration/stability properties imply different operational uses.
  • Research agenda:
    • Further work should formalize the proposed adaptability–stability–interpretability framework, study intervention strategies when models disagree, and explore techniques to reduce preprocessing sensitivity in highly adaptive architectures.

Assessment

Paper Typecorrelational Evidence Strengthmedium — The paper uses rigorous predictive-evaluation practices (out-of-sample splits across a clear regime break, cross-validation, bootstrap confidence intervals, and multi-country data), which supports reliable claims about relative model performance and calibration; however, it is not designed to identify causal effects and the results may be sensitive to sample selection, preprocessing, and model/hyperparameter choices. Methods Rigormedium — Evaluation uses appropriate predictive-validation techniques (pre/post regime splits, CV, bootstrap CIs) and compares multiple baselines, but important details are missing or raise potential concerns (extent of hyperparameter tuning, exact preprocessing pipelines, how bankruptcy is defined/labelled, handling of class imbalance beyond 'balanced panel', and sensitivity analyses to alternative sample definitions), and interpretability relies on SHAP whose magnitudes are not directly comparable across models. SampleBalanced panel of publicly listed firms in the United States and China covering 2015–2022, predicting 1-year-ahead bankruptcy/distress; samples are split pre- and post-COVID for out-of-sample evaluation; exact number of firms/observations, treatment of delisted firms, and class imbalance statistics are not provided in the supplied text. Themesgovernance adoption GeneralizabilityOnly publicly listed firms in the U.S. and China — excludes private firms and SMEs, Time period limited to 2015–2022; results may not hold outside these years or for other macro shocks, Balanced-panel construction may induce selection bias (survivor bias) that limits applicability to broader populations, Model performance and calibration could differ with other preprocessing pipelines, feature sets, or hyperparameter tuning, SHAP attributions are only interpreted within-model and are not comparable across architectures or institutional contexts, Findings on bankruptcy prediction may not generalize to other financial outcomes or to non-financial prediction tasks

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
NGBoost–SHAP produces more stable 1-year bankruptcy probabilities than TabNet–SHAP across the U.S. and Chinese markets and across the COVID regime break. Decision Quality positive Stability of predicted 1-year bankruptcy probabilities across markets and economic regimes
Reading fidelity high
Study strength medium
not reported
0.3
NGBoost–SHAP provides better probabilistic calibration of 1-year bankruptcy predictions than TabNet–SHAP. Decision Quality positive Calibration of predicted 1-year bankruptcy probabilities
Reading fidelity high
Study strength medium
not reported
0.3
TabNet–SHAP exhibits greater variability and higher sensitivity to regime shifts and preprocessing choices than NGBoost–SHAP. Decision Quality negative Prediction variability and sensitivity to economic-regime changes and preprocessing choices
Reading fidelity high
Study strength medium
not reported
0.3
The difference between NGBoost–SHAP and TabNet–SHAP reflects an architecture-specific trade-off between stability and adaptability rather than an absolute ranking of the models. Decision Quality mixed Relative stability and adaptability of bankruptcy-prediction models across regimes
Reading fidelity high
Study strength medium
not reported
0.3
Profitability and capital-structure variables carry most of the explanatory signal in both markets and both models. Decision Quality positive Within-model feature attribution for 1-year bankruptcy predictions
Reading fidelity high
Study strength medium
not reported
0.3
The post-COVID Chinese subsample has the most concentrated feature-attribution pattern. Decision Quality positive Concentration of SHAP feature attributions
Reading fidelity high
Study strength medium
not reported
0.3
SHAP values should be interpreted as within-model feature rankings and are not directly comparable across models or markets. Ai Safety And Ethics null_result Cross-model and cross-market comparability of feature-attribution magnitudes
Reading fidelity high
Study strength low
not reported
0.15

Notes