predictive AI models that estimate probabilities for a binary outcome to support medical decisions
Source article: Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models intended for use in medical practice. Poorly performing models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs. In this Viewpoint, we assess the merits of classic and contemporary performance measures when validating predictive AI models for med…
Contested: both sides are scored from claims and sources, not community votes.

"Friends of Cancer Research" by MDGovpics is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
Published in December 2025, this Viewpoint in The Lancet Digital Health evaluates how to validate predictive AI models that estimate binary outcome probabilities for clinical use. The authors reviewed 32 performance measures across five domains and examined whether each measure is proper and whether it accounts for misclassification costs, illustrating findings with the ADNEX model for ovarian tumour malignancy.
The work matters because improper or incomplete validation can lead to misleading models and harmful clinical decisions, while proper evaluation supports safer adoption. Uncertainty remains about how these recommendations will be implemented across diverse clinical settings and whether reporting of AUROC, calibration plots, and decision curve analysis will become standard practice for binary medical AI models.
- Assessed 32 performance measures across five domains: discrimination, calibration, overall performance, classification, and clinical utility.
- Evaluated two key characteristics: whether expected value is optimised using correct probabilities (proper measure) and whether measure accounts for misclassification costs.
- Found 17 measures showed both characteristics, 14 showed one, and one (F1 score) showed neither.
- Illustrated measures using the ADNEX model which predicts probability of malignancy in women with an ovarian tumour.
Appropriate evaluation using proper measures including AUROC, calibration plot, and net benefit with decision curve analysis is essential to validate predictive AI models that estimate binary outcome probabilities for medical practice.
Poorly performing predictive AI models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs, with classification measures being improper at clinically relevant thresholds.
The rundown
The Viewpoint systematically reviewed 32 measures and graphical assessments across discrimination, calibration, overall performance, classification, and clinical utility, distinguishing statistical performance from decision-analytical performance.
Using the ADNEX ovarian tumour malignancy model as an illustration, authors found only 17 measures were both proper and accounted for misclassification costs, while F1 score was neither, and all classification measures were improper except at threshold 05 or true prevalence.
Guidance is scoped to models that estimate probabilities for a binary outcome, not multiclass, continuous, or other prediction tasks.
Sources
- Peer-reviewedThe Lancet Digital Health2025-12-01
How should this claim be treated?
ace
The debate