comparative discrimination of machine learning versus logistic regression for trauma mortality prediction on same datasets
Source article: Systematic Bias in Comparative Evaluations of Machine Learning Versus Logistic Regression for Clinical Prediction Models: A Meta-Research Analysis Using Trauma Mortality as an Empirical Case
Abstract: Objective Comparative evaluations of machine learning (ML) and logistic regression (LR) for clinical prediction frequently report ML as superior, but the methodological framework producing those comparisons has received limited scrutiny. We aimed to quantify the apparent discrimination advantage of ML over LR using trauma mortality prediction as an empirical case, and to characterise the evaluation practices that shape it. Study design and setting Systematic review and random-effects meta-analysis combined with…
Contested: both sides are scored from claims and sources, not community votes.
Conditions pour IA digne de confiance conditions requirements for Trustworthy AI by Lamiot. CC BY 4.0 · https://creativecommons.org/licenses/by/4.0
A systematic review and meta-research appraisal examined 20 studies comparing machine learning and logistic regression for trauma mortality prediction, with 17 studies (243,324 patients) in primary synthesis. The pooled within-study AUC difference favoring the best ML model was 0.026 (95% CI 0.009-0.043), 0.017 in co-primary analysis of studies reporting CIs, with extreme heterogeneity and a prediction interval crossing zero.
The analysis matters because claims that ML outperforms traditional models influence clinical adoption, yet the observed edge was small, inconsistent, and potentially inflated by design choices such as best-of-tournament ML versus single LR, internal validation, and AUC-only reporting. Uncertainty remains about true comparative performance under fair, externally validated, and fully reported evaluations including calibration and decision analysis.
- Systematic review and random-effects meta-analysis of studies directly comparing any ML algorithm with LR on same dataset for trauma mortality.
- Co-primary analysis restricted to studies reporting confidence intervals for both models yielded pooled difference 0.017 (95% CI 0.005-0.029).
- Advantage larger for best-of-tournament ensemble methods (0.034) than single ML algorithms (0.005), interaction not statistically significant.
- Only three studies used external or temporal validation; four of 17 achieved low PROBAST risk of bias.
Across 17 studies totaling 243,324 trauma patients, the best-performing ML model showed a small pooled AUC advantage over logistic regression for mortality prediction.
Apparent superiority of ML over logistic regression for trauma mortality prediction may be inflated by convergent practices including comparing best-of-several ML models to a single LR comparator, reliance on internal validation, selective reporting, and AUC-only synthesis.
The rundown
Search covered MEDLINE (Ovid), Scopus, Web of Science, and Embase through 15 January 2025 for direct within-dataset comparisons, PROSPERO CRD42025636303, with sensitivity search during revision. Estimand was pre-specified as within-study AUC difference between best-performing ML and single LR comparator, noted as itself a source of bias.
Authors propose six minimum standards: pre-specification, fair comparator design, robust external or temporal validation, calibration and decision-analytic reporting, full transparency, and bias-aware synthesis, arguing these may be relevant beyond trauma to other clinical prediction settings.
Extreme between-study heterogeneity and wide prediction interval mean future studies could favor either approach, and most included studies had high or unclear risk of bias with internal validation only.
Sources
- Peer-reviewedJournal of Clinical Epidemiology2026-09-11
How should this claim be treated?
ace
The debate