TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-10

machine-learning prediction of urine-culture positivity from routine urinalysis data in patients with paired orders

Source article: Machine-learning prediction of urine-culture positivity in a multicentre test-ordered cohort: Model development and internal validation

Abstract: Objectives This study aim to develop, compare and internally validate machine-learning models for predicting urine-culture positivity in patients who had both urinalysis and culture ordered and to explore descriptive probability strata. Post hoc secondary analyses examined age subgroups, the incremental contribution of text-derived features, simpler comparators and calibration. Patients and methods Urine culture results are typically unavailable for 24-72 h, creating uncertainty during initial assessment, and ma…

TRV-2026-1053Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 70The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 68The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Machine-learning prediction of urine-culture positivity in a multicentre test-ordered cohort: Model development and internal validation

United States Naval Medical Bulletin Vol. 42, Nos. 1-6, 1944 by U.S. Navy. Bureau of Medicine and Surgery. Public domain

The quick read

Researchers developed and internally validated machine-learning models to estimate the probability of urine-culture positivity using routinely collected urinalysis data from 2530 sample records across three university hospitals. Using a stratified 75:25 sample-level split, 13 supervised algorithms were tested, with CatBoost showing the highest test-set AUC of 0.858 (95% CI 0.829-0.892) and similar performance to other gradient-boosting models.

The result matters because culture results take 24-72 hours, leaving uncertainty during initial assessment, and a risk-estimation tool could eventually help prioritize or interpret testing. As of the September 2026 publication date, the work remains internal validation only, without patient-grouped or external validation, without linkage to symptomatic UTI or antibiotic outcomes, and with explicit statements that clinical utility and safety were not evaluated.

Main points
  • Retrospective study of 2530 urine-sample records from three university hospitals where both urinalysis and culture were ordered.
  • Thirteen supervised algorithms evaluated using stratified 75:25 sample-record split for internal validation.
  • CatBoost numerically highest AUC 0.858 (95% CI 0.829-0.892), not significantly different from gradient boosting or XGBoost.
  • At operating threshold sensitivity 0.587, specificity 0.930, PPV 0.766, NPV 0.851.
  • Exploratory probability strata separated different observed positivity rates but clinical utility and safety not evaluated.
Gain

In 2530 paired urinalysis-culture records from three university hospitals, gradient-boosting models estimated culture positivity after urinalysis, with CatBoost achieving test-set AUC 0.858 and high specificity at the reported threshold.

Problem

Models were validated only at sample level without patient or centre grouping, and exploratory risk strata were not evaluated for clinical utility or safety, so they do not establish symptomatic UTI or safe antibiotic decisions.

The rundown

The study used 2530 retrospective urine-sample records from three university hospitals, including only cases where both urinalysis and culture were ordered, not symptom-based UTI criteria. Thirteen algorithms were compared on a 75:25 split, with gradient-boosting methods showing similar discrimination.

CatBoost reached AUC 0.858 (95% CI 0.829-0.892) on the test set, with sensitivity 0.587, specificity 0.930, PPV 0.766 and NPV 0.851 at the reported threshold. Authors noted records were not grouped by patient or centre due to unavailable stable identifiers, and post hoc analyses looked at age subgroups, text-derived features and calibration.

What this doesn’t fix

Internal sample-level validation only without patient or centre grouping, culture positivity not equivalent to symptomatic UTI, and no evaluation of clinical outcomes or antibiotic decision safety; external validation required before clinical use.

Sources

Reader signal

How should this claim be treated?

The debate