An Interpretable Machine Learning Framework with Clinical Nomogram for Predicting In-Hospital Mortality in Acute Ischemic Stroke Using High-Granularity Bedside Data
Abstract: This multicenter study developed and validated an interpretable machine learning model integrating granular nursing and emergency department data collected within the first 24 hours to predict in-hospital mortality in acute ischemic stroke (AIS). We analyzed a retrospective cohort of 5,014 adult AIS patients from three tertiary academic centers (2019-2023). Centers A and B (n=3,512) formed the development cohort; Center C (n=1,502) served as the external validation cohort. Sixty-three predictors across seven dom…

"Brain trauma CT" by Rehman T, Ali R, Tawil I, Yonas H is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
Researchers developed and externally validated an interpretable machine learning framework to predict in-hospital mortality in acute ischemic stroke using high-granularity bedside data from the first 24 hours. Using 5,014 patients from three tertiary centers between 2019-2023, the best CatBoost model with 22 features achieved AUC-ROC 0.917 internally and 0.891 externally, outperforming traditional ICU scores.
The result matters because early risk stratification in stroke relies on generic ICU scores that may miss nursing assessment signals, and the study shows ensemble gradient boosting with SHAP and a nomogram can improve discrimination and calibration. Uncertainty remains about prospective deployment, workflow integration, and generalizability beyond the three academic centers studied.
- Retrospective cohort of 5,014 adult AIS patients from three tertiary academic centers (2019-2023), with Centers A and B (n=3,512) for development and Center C (n=1,502) for external validation.
- Sixty-three predictors across seven domains from initial 24 hours; consensus feature selection combining LASSO, RFE-RF, filter methods, and XGBoost importance; six algorithms evaluated with nested cross-validation and Bayesian optimization.
- Best CatBoost model with 22 RFE-RF features had calibration slopes 0.934-0.977 and Brier 0.098-0.114; 0-12h sensitivity analysis showed external AUC 0.873.
- 10-feature logistic nomogram achieved external AUC 0.871 with superior DCA net benefit over comparators, with SHAP values used for interpretability.
A CatBoost model integrating 22 granular nursing and emergency department features collected within the first 24 hours improved early in-hospital mortality prediction for acute ischemic stroke patients, achieving higher discrimination than established ICU scores in both internal and external validation.
The rundown
The study extracted 63 predictors across seven domains from the first 24 hours of nursing and emergency department data and applied anti-leakage protocols, SMOTE, and nested cross-validation to compare six machine learning algorithms.
The selected CatBoost model used 22 RFE-RF features and was compared to APACHE III, SOFA, OASIS, and GCS with reported delta AUC 0.142-0.193, and a 10-feature nomogram was also validated for clinical use.
Sources
- Peer-reviewedJournal of Stroke and Cerebrovascular Diseases2026-08-26
How should this claim be treated?
ace
The debate