machine learning prediction of bacterial infections in infants aged 1 to 90 days using routine non-invasive paraclinical markers versus CSF culture gold standard

Source article: Machine learning models for predicting neonatal bacterial infections: a retrospective cohort study

Bacterial infections represent a critical threat to neonatal health, accounting for approximately 25% of neonatal mortality globally. Timely and precise diagnosis in infants aged 1 to 90 days is essential to facilitate rapid intervention and prevent severe complications. This study aimed to develop and evaluate machine learning (ML) models for the early, non-invasive prediction of bacterial infections using routine clinical data, maximizing clinical interpretability for point-of-care triage. Data from 306 infant…

Machine learning models for predicting neonatal bacterial infections: a retrospective cohort study
Hospital Universitari Doctor Peset, València 06 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
Trace impact readingPositive state
P 68The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.

Both sides are scored from claims and sources, not community votes.

G 74The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.

In brief

Researchers retrospectively analyzed 306 infants aged 1 to 90 days hospitalized between 2014 and 2022 in Khorasan Razavi, Iran, using CSF culture via lumbar puncture as the gold standard, to train nine machine learning classifiers on routine non-invasive paraclinical markers with nested cross-validation and SHAP interpretation.

The work matters because bacterial infections account for about 25% of neonatal mortality globally and lumbar puncture is invasive; the models offer an objective risk-stratification aid at point-of-care triage, but the documented sensitivity-specificity trade-off and optimism gaps mean it remains uncertain whether the tool can reduce unnecessary procedures without missing infections in broader populations.

Main points

  1. Retrospective analysis of 306 infants aged 1 to 90 days hospitalized 2014-2022 with CSF culture via lumbar puncture as gold standard (158 infectious, 148 non-infectious).
  2. Nine classifiers evaluated via leakage-safe nested cross-validation (5-folds x 2 repeats outer, threefold inner) with QuantileTransformer normalization and SHAP interpretation.
  3. HistGBM had highest raw AUROC 0.786 (95% CI: 0.757-0.812) but severe training optimism gap 0.214; L2-regularized logistic regression selected for minimal gap 0.083 and Brier score 0.208.

The gain

L2-regularized logistic regression trained on routine non-invasive paraclinical markers achieved stable discrimination with minimal generalization gap and robust calibration for early prediction of bacterial infections in infants aged 1 to 90 days.

The problem

When tuned for ≥95% sensitivity screening, the models suffered a steep parallel decline in specificity and cannot safely eliminate the need for lumbar puncture, with top non-linear models showing severe training optimism.

The rundown

The study used 306 infants from a Social Security Organization hospital in Khorasan Razavi, Iran, labeled by CSF culture via lumbar puncture, and limited predictors to routine non-invasive paraclinical markers from the EHR normalized with QuantileTransformer.

Evaluation used a leakage-safe nested cross-validation framework and SHAP values for global and local interpretability, with training-to-validation audits showing top-tier models clustered at outer-CV AUROC 0.74-0.76 and predictive signal distributed across urinary markers, age, and metabolic indicators.

What this doesn’t fix

High-sensitivity operation forces steep specificity loss, so model cannot be used as standalone rule-out to eliminate lumbar punctures and must remain a risk-tiering aid.

Sources

  1. Peer-reviewedEuropean Journal of Pediatrics2026-08-14

The debate