Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models intended for use in medical practice. Poorly performing models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs. In this Viewpoint, we assess the merits of classic and contemporary performance measures when validating predictive AI models for med…
Appropriate evaluation using proper measures including AUROC, calibration plot, and net benefit with decision curve analysis is essential to validate predictive AI models that estimate binary outcome probabilities for medical practice.
Poorly performing predictive AI models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs, with classification measures being improper at clinically relevant thresholds.
Guidance is scoped to models that estimate probabilities for a binary outcome, not multiclass, continuous, or other prediction tasks.
Evidence
- Peer-reviewedThe Lancet Digital Health2025-12-01
How should this claim be treated?
Truvace Impact Record TRV-2026-0434, v1: “Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance.” Truvace, 2026-07-20. /record/TRV-2026-0434 (accessed at citation time). sha256 3185e1606a622eeb…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0434 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace