0 cumulative citations
View corpus contextProbabilistic boosting outperforms deep tabular nets for stable credit-risk probabilities: NGBoost–SHAP gives better-calibrated, less regime-sensitive 1‑year bankruptcy forecasts across US and China, whereas TabNet–SHAP adapts faster but is more variable and preprocessing-sensitive.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextABSTRACT Despite substantial advances in machine‐learning approaches to bankruptcy prediction, less is known about whether explainable models remain reliable across markets and during structural breaks. This study examines that issue by comparing two explainable architectures, TabNet–SHAP and NGBoost–SHAP, for 1‐year‐ahead bankruptcy prediction in the United States and China during 2015–2022. Using a balanced panel of listed firms from both markets, we evaluate out‐of‐sample performance before and after COVID‐19 and benchmark both models against conventional linear and ensemble classifiers. Model uncertainty is assessed through cross‐validation and bootstrap confidence intervals. The results show that NGBoost–SHAP provides more stable and better‐calibrated probabilistic predictions across regimes, whereas TabNet–SHAP produces more variable results and is more sensitive to regime shifts and preprocessing choices. We interpret this pattern as an architecture‐specific trade‐off between stability and adaptability rather than as an unconditional ranking of the two models. SHAP‐based interpretation indicates that profitability and capital‐structure variables account for most of the explanatory signal within each market and model. The post‐COVID Chinese subsample shows the most concentrated feature‐attribution structure, although SHAP magnitudes are interpreted only as within‐model rankings rather than as directly comparable cross‐market quantities. The study contributes by proposing an adaptability–stability–interpretability framework for selecting bankruptcy prediction models under different institutional and market conditions. It also provides evidence that explainable ensemble and deep tabular models can recover economically meaningful distress signals across markets while differing in their calibration, stability, and sensitivity to economic regime change.
Summary
Main Finding
NGBoost–SHAP yields more stable and better‑calibrated 1‑year bankruptcy probabilities across the U.S. and Chinese markets and across the COVID regime break, whereas TabNet–SHAP is more adaptive but produces more variable, regime‑sensitive results. This pattern reflects an architecture‑specific trade‑off between stability and adaptability rather than an absolute ranking of the two models.
Key Points
- Models compared: TabNet–SHAP (deep tabular architecture with SHAP explanations) and NGBoost–SHAP (probabilistic boosting with SHAP).
- Scope: 1‑year‑ahead bankruptcy prediction for listed firms in the United States and China, 2015–2022, using a balanced panel.
- Evaluation: out‑of‑sample performance before and after the COVID shock, benchmarked against conventional linear and ensemble classifiers.
- Uncertainty quantification: cross‑validation and bootstrap confidence intervals used to assess model uncertainty and stability.
- Main empirical results:
- NGBoost–SHAP produced more stable predictions across regimes and better probabilistic calibration.
- TabNet–SHAP showed greater variability, higher sensitivity to regime shifts and preprocessing choices, and less stable calibration.
- Explainability findings (SHAP):
- Profitability and capital‑structure variables carry most of the explanatory signal in both markets and models.
- The post‑COVID Chinese subsample exhibited the most concentrated feature‑attribution pattern.
- SHAP values are treated as within‑model feature rankings and are not directly comparable across models or markets.
Data & Methods
- Data: Balanced panel of publicly listed firms in the U.S. and China covering 2015–2022.
- Prediction target: 1‑year ahead bankruptcy/distress.
- Models evaluated:
- TabNet with SHAP for local/global feature attributions.
- NGBoost with SHAP to provide probabilistic forecasts and feature attributions.
- Conventional linear and ensemble baselines for benchmarking.
- Validation strategy:
- Out‑of‑sample tests split around the COVID regime break (pre/post).
- Cross‑validation for model selection and stability checks.
- Bootstrap procedures to compute confidence intervals for performance metrics and SHAP summaries.
- Interpretability: SHAP used to rank feature importance within each model; magnitudes interpreted cautiously as relative (within‑model) signals.
Implications for AI Economics
- Model selection under regime change:
- Use NGBoost‑like probabilistic ensembles when stability and well‑calibrated risk probabilities are critical (e.g., regulatory stress tests, credit scoring under varying macro states).
- Consider TabNet‑like deep tabular models for environments where adaptability to new patterns may pay off, but pair them with rigorous validation and robustness checks.
- Model monitoring and governance:
- Regular out‑of‑sample testing across economic regimes and bootstrap uncertainty quantification are essential to detect calibration drift and preprocessing sensitivity.
- SHAP summaries are useful to recover economically meaningful drivers (profitability, leverage) but should not be used to compare importance magnitudes across models or markets without caution.
- Policy and risk management:
- Explainable ML can recover relevant distress signals across institutional contexts, supporting cross‑market credit surveillance, but differing calibration/stability properties imply different operational uses.
- Research agenda:
- Further work should formalize the proposed adaptability–stability–interpretability framework, study intervention strategies when models disagree, and explore techniques to reduce preprocessing sensitivity in highly adaptive architectures.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| NGBoost–SHAP produces more stable 1-year bankruptcy probabilities than TabNet–SHAP across the U.S. and Chinese markets and across the COVID regime break. Decision Quality | positive | Stability of predicted 1-year bankruptcy probabilities across markets and economic regimes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| NGBoost–SHAP provides better probabilistic calibration of 1-year bankruptcy predictions than TabNet–SHAP. Decision Quality | positive | Calibration of predicted 1-year bankruptcy probabilities |
Reading fidelity
high
Study strength
medium
|
not reported
|
| TabNet–SHAP exhibits greater variability and higher sensitivity to regime shifts and preprocessing choices than NGBoost–SHAP. Decision Quality | negative | Prediction variability and sensitivity to economic-regime changes and preprocessing choices |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The difference between NGBoost–SHAP and TabNet–SHAP reflects an architecture-specific trade-off between stability and adaptability rather than an absolute ranking of the models. Decision Quality | mixed | Relative stability and adaptability of bankruptcy-prediction models across regimes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Profitability and capital-structure variables carry most of the explanatory signal in both markets and both models. Decision Quality | positive | Within-model feature attribution for 1-year bankruptcy predictions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The post-COVID Chinese subsample has the most concentrated feature-attribution pattern. Decision Quality | positive | Concentration of SHAP feature attributions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| SHAP values should be interpreted as within-model feature rankings and are not directly comparable across models or markets. Ai Safety And Ethics | null_result | Cross-model and cross-market comparability of feature-attribution magnitudes |
Reading fidelity
high
Study strength
low
|
not reported
|