TruaceTracing the truth around AITuesday, July 21, 2026
Health·The Trace·Automated dual reading·Published 2026-07-20

predictive AI models that estimate probabilities for a binary outcome to support medical decisions

Source article: Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance

Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models intended for use in medical practice. Poorly performing models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs. In this Viewpoint, we assess the merits of classic and contemporary performance measures when validating predictive AI models for med…

TRV-2026-0434Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 75The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 72The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance

"Friends of Cancer Research" by MDGovpics is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.

The quick read

Published in December 2025, this Viewpoint in The Lancet Digital Health evaluates how to validate predictive AI models that estimate binary outcome probabilities for clinical use. The authors reviewed 32 performance measures across five domains and examined whether each measure is proper and whether it accounts for misclassification costs, illustrating findings with the ADNEX model for ovarian tumour malignancy.

The work matters because improper or incomplete validation can lead to misleading models and harmful clinical decisions, while proper evaluation supports safer adoption. Uncertainty remains about how these recommendations will be implemented across diverse clinical settings and whether reporting of AUROC, calibration plots, and decision curve analysis will become standard practice for binary medical AI models.

Main points
  • Assessed 32 performance measures across five domains: discrimination, calibration, overall performance, classification, and clinical utility.
  • Evaluated two key characteristics: whether expected value is optimised using correct probabilities (proper measure) and whether measure accounts for misclassification costs.
  • Found 17 measures showed both characteristics, 14 showed one, and one (F1 score) showed neither.
  • Illustrated measures using the ADNEX model which predicts probability of malignancy in women with an ovarian tumour.
Gain

Appropriate evaluation using proper measures including AUROC, calibration plot, and net benefit with decision curve analysis is essential to validate predictive AI models that estimate binary outcome probabilities for medical practice.

Problem

Poorly performing predictive AI models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs, with classification measures being improper at clinically relevant thresholds.

The rundown

The Viewpoint systematically reviewed 32 measures and graphical assessments across discrimination, calibration, overall performance, classification, and clinical utility, distinguishing statistical performance from decision-analytical performance.

Using the ADNEX ovarian tumour malignancy model as an illustration, authors found only 17 measures were both proper and accounted for misclassification costs, while F1 score was neither, and all classification measures were improper except at threshold 05 or true prevalence.

What this doesn’t fix

Guidance is scoped to models that estimate probabilities for a binary outcome, not multiclass, continuous, or other prediction tasks.

Sources

Reader signal

How should this claim be treated?

The debate