TruaceTracing the truth around AIMonday, September 14, 2026
TRV-2026-1068Version 1 · Certified

Written 2026-09-13 06:56:15 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1068
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-13T06:56:15.503259Z
status: published
lens: trace
sector: health
headline: Systematic Bias in Comparative Evaluations of Machine Learning Versus Logistic Regression for Clinical Prediction Models: A Meta-Research Analysis Using Trauma Mortality as an Empirical Case
dek: Objective Comparative evaluations of machine learning (ML) and logistic regression (LR) for clinical prediction frequently report ML as superior, but the methodological framework producing those comparisons has received limited scrutiny. We aimed to quantify the apparent discrimination advantage of ML over LR using trauma mortality prediction as an empirical case, and to characterise the evaluation practices that shape it. Study design and setting Systematic review and random-effects meta-analysis combined with…
gain_title: Across 17 studies totaling 243,324 trauma patients, the best-performing ML model showed a small pooled AUC advantage over logistic regression for mortality prediction.
problem_title: Apparent superiority of ML over logistic regression for trauma mortality prediction may be inflated by convergent practices including comparing best-of-several ML models to a single LR comparator, reliance on internal validation, selective reporting, and AUC-only synthesis.
trace_subject: comparative discrimination of machine learning versus logistic regression for trauma mortality prediction on same datasets
gain_reading: Across 17 studies totaling 243,324 trauma patients, the best-performing ML model showed a small pooled AUC advantage over logistic regression for mortality prediction.
gain_evidence: The pooled AUC difference favouring ML was 0.026 (95% CI 0.009-0.043) | 17 (243,324 patients) contributed to the primary quantitative synthesis
problem_reading: Apparent superiority of ML over logistic regression for trauma mortality prediction may be inflated by convergent practices including comparing best-of-several ML models to a single LR comparator, reliance on internal validation, selective reporting, and AUC-only synthesis.
problem_evidence: Four convergent evaluation practices - model-selection asymmetry, reliance on internal validation, selective reporting, and AUC-only synthesis - may jointly inflate apparent ML superiority and are not addressed by current evidence-synthesis frameworks
quick_read: A systematic review and meta-research appraisal examined 20 studies comparing machine learning and logistic regression for trauma mortality prediction, with 17 studies (243,324 patients) in primary synthesis. The pooled within-study AUC difference favoring the best ML model was 0.026 (95% CI 0.009-0.043), 0.017 in co-primary analysis of studies reporting CIs, with extreme heterogeneity and a prediction interval crossing zero.

The analysis matters because claims that ML outperforms traditional models influence clinical adoption, yet the observed edge was small, inconsistent, and potentially inflated by design choices such as best-of-tournament ML versus single LR, internal validation, and AUC-only reporting. Uncertainty remains about true comparative performance under fair, externally validated, and fully reported evaluations including calibration and decision analysis.
limitation: Extreme between-study heterogeneity and wide prediction interval mean future studies could favor either approach, and most included studies had high or unclear risk of bias with internal validation only.
tag: Dual reading
key_points: Systematic review and random-effects meta-analysis of studies directly comparing any ML algorithm with LR on same dataset for trauma mortality. | Co-primary analysis restricted to studies reporting confidence intervals for both models yielded pooled difference 0.017 (95% CI 0.005-0.029). | Advantage larger for best-of-tournament ensemble methods (0.034) than single ML algorithms (0.005), interaction not statistically significant. | Only three studies used external or temporal validation; four of 17 achieved low PROBAST risk of bias.
rundown: Search covered MEDLINE (Ovid), Scopus, Web of Science, and Embase through 15 January 2025 for direct within-dataset comparisons, PROSPERO CRD42025636303, with sensitivity search during revision. Estimand was pre-specified as within-study AUC difference between best-performing ML and single LR comparator, noted as itself a source of bias.

Authors propose six minimum standards: pre-specification, fair comparator design, robust external or temporal validation, calibration and decision-analytic reporting, full transparency, and bias-aware synthesis, arguing these may be relevant beyond trauma to other clinical prediction settings.
sources:
- peer_reviewed | Journal of Clinical Epidemiology | https://doi.org/10.1016/j.jclinepi.2026.112508 | 2026-09-11
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
16afba5e89a1c3173aeb8456478aa7cd29cc17eb497448c9b71443093148333c
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1068 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.