Using interpretable machine learning models to predict the occurrence of severe influenza in hospitalized children: a retrospective cohort study

Influenza, a prevalent disease, significantly threatens public health. Accurately predicting severe influenza occurrences is crucial for developing personalized prevention strategies and treatment plans. This study aimed to construct a highly interpretable model to assess the risk of severe influenza in hospitalized children, using the SHapley Additive exPlanation (SHAP) method to interpret the Random Forest (RF) model and identify risk factors for severe influenza. A retrospective cohort study was conducted, co…

Using interpretable machine learning models to predict the occurrence of severe influenza in hospitalized children: a retrospective cohort study
Hospital Universitari Doctor Peset, València 11 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

In brief

Researchers conducted a retrospective cohort study of 436 hospitalized children with influenza, using electronic medical record data from the first 24 hours after admission to train and test six machine learning models. The Random Forest model achieved the highest performance with an AUROC of 0.88 and greater net benefit on decision curve analysis, with SHAP analysis highlighting blood glucose as the most significant predictive variable.

The result matters because early identification of children at risk for severe influenza could enable personalized prevention and treatment planning in pediatric inpatient care. What remains uncertain from this text is external generalizability beyond the single EMR source, prospective performance, and whether the identified predictors lead to changed clinical actions or outcomes.

Main points

  1. Retrospective cohort of 436 eligible influenza patients from EMR system, split 70% training and 30% verification.
  2. Data collected within the first 24 h after patient admission for model development.
  3. SHAP method used to interpret Random Forest model and rank risk factors.

The gain

A Random Forest model trained on first-24-hour EMR data predicted severe influenza in hospitalized children with AUROC 0.88, outperforming five other models on decision curve analysis, with blood glucose ranked as top predictor.

The rundown

The study collected hospitalization records of influenza patients from the Electronic Medical Records (EMR) system and used data from the first 24 hours after admission. The dataset was randomly divided into 70% for training and 30% for accuracy verification across six compared models.

Interpretability was provided by the SHapley Additive exPlanation method applied to the Random Forest model, which also showed higher net benefit on decision curve analysis than other machine learning models.

Sources

  1. Peer-reviewedBMC Medical Informatics and Decision Making2026-10-05

The debate