TruaceTracing the truth around AIWednesday, July 22, 2026
TRV-2026-0434Version 1 · Certified

Written 2026-07-20 10:47:43 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0434
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-20T10:47:43.268467Z
status: published
lens: trace
sector: health
headline: Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
dek: Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models intended for use in medical practice. Poorly performing models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs. In this Viewpoint, we assess the merits of classic and contemporary performance measures when validating predictive AI models for med…
gain_title: Appropriate evaluation using proper measures including AUROC, calibration plot, and net benefit with decision curve analysis is essential to validate predictive AI models that estimate binary outcome probabilities for medical practice.
problem_title: Poorly performing predictive AI models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs, with classification measures being improper at clinically relevant thresholds.
trace_subject: predictive AI models that estimate probabilities for a binary outcome to support medical decisions
gain_reading: Appropriate evaluation using proper measures including AUROC, calibration plot, and net benefit with decision curve analysis is essential to validate predictive AI models that estimate binary outcome probabilities for medical practice.
gain_evidence: Selecting appropriate performance measures is essential for predictive AI models intended for use in medical practice. | We recommend the following measures and plots as essential to report: area under the receiver operating characteristic curve, calibration plot, a clinical utility measure such as net benefit with decision curve analysis
problem_reading: Poorly performing predictive AI models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs, with classification measures being improper at clinically relevant thresholds.
problem_evidence: Poorly performing models are misleading and may lead to wrong clinical decisions that can be detrimental to patients and increase financial costs. | All classification measures were improper for clinically relevant decision thresholds other than when the threshold was 05 or equal to the true prevalence.
quick_read: Published in December 2025, this Viewpoint in The Lancet Digital Health evaluates how to validate predictive AI models that estimate binary outcome probabilities for clinical use. The authors reviewed 32 performance measures across five domains and examined whether each measure is proper and whether it accounts for misclassification costs, illustrating findings with the ADNEX model for ovarian tumour malignancy.

The work matters because improper or incomplete validation can lead to misleading models and harmful clinical decisions, while proper evaluation supports safer adoption. Uncertainty remains about how these recommendations will be implemented across diverse clinical settings and whether reporting of AUROC, calibration plots, and decision curve analysis will become standard practice for binary medical AI models.
limitation: Guidance is scoped to models that estimate probabilities for a binary outcome, not multiclass, continuous, or other prediction tasks.
tag: Automated dual reading
key_points: Assessed 32 performance measures across five domains: discrimination, calibration, overall performance, classification, and clinical utility. | Evaluated two key characteristics: whether expected value is optimised using correct probabilities (proper measure) and whether measure accounts for misclassification costs. | Found 17 measures showed both characteristics, 14 showed one, and one (F1 score) showed neither. | Illustrated measures using the ADNEX model which predicts probability of malignancy in women with an ovarian tumour.
rundown: The Viewpoint systematically reviewed 32 measures and graphical assessments across discrimination, calibration, overall performance, classification, and clinical utility, distinguishing statistical performance from decision-analytical performance.

Using the ADNEX ovarian tumour malignancy model as an illustration, authors found only 17 measures were both proper and accounted for misclassification costs, while F1 score was neither, and all classification measures were improper except at threshold 05 or true prevalence.
sources:
- peer_reviewed | The Lancet Digital Health | https://doi.org/10.1016/j.landig.2025.100916 | 2025-12-01
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
3185e1606a622eeb6a597af7597e27f0dcaee276eab43e2f6f8c15a5dc9354e5
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0434 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.