diagnostic accuracy of generative AI for radiology tasks

Source article: Generative AI versus physicians in diagnostic radiology: a systematic review and meta-analysis

Purpose Generative artificial intelligence (AI) models are increasingly evaluated for diagnostic tasks in radiology, yet accuracy, study designs, endpoints, and comparators vary widely. The purpose was to synthesize diagnostic accuracy of generative AI for radiology and compare performance with physicians. Materials and methods A systematic review and meta-analysis was prospectively registered in PROSPERO (CRD420251040000) and conducted in accordance with PRISMA-DTA guidance. Searches of Medline, Scopus, Web of…

Generative AI versus physicians in diagnostic radiology: a systematic review and meta-analysis
Hospital Universitari Doctor Peset, València 09 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
Trace impact readingContested
P 73The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.

Both sides are scored from claims and sources, not community votes.

G 71The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.

In brief

A systematic review and meta-analysis of 48 studies published through March 2025 synthesized diagnostic accuracy of generative AI in radiology and compared it to physician performance using multilevel random-effects meta-regression.

The finding matters because lower AI accuracy relative to experts raises safety concerns for clinical deployment, while the higher accuracy with text-only inputs suggests assistive, text-oriented use cases need further study but may be confounded by task difficulty and information content.

Main points

  1. Systematic review and meta-analysis of 48 studies from June 2018-March 2025 comparing generative AI to physicians on radiology diagnostic tasks.
  2. Pooled generative AI accuracy was 42.9% for free-text tasks and 58.1% for choice tasks.
  3. Multilevel random-effects meta-regression with study-clustered robust inference found text-only input outperformed image-only and text-and-image inputs.
  4. Authors concluded standardized, transparently reported, adequately powered evaluations are warranted before clinical deployment.

The gain

Generative AI achieved higher diagnostic accuracy when given text-only input compared to image-only input in radiology tasks.

The problem

Generative AI showed significantly lower diagnostic accuracy than expert physicians on radiology tasks, with a 13.0 percentage point gap.

The rundown

The review was prospectively registered in PROSPERO and followed PRISMA-DTA guidance, searching Medline, Scopus, Web of Science, Cochrane Central, and medRxiv, with two reviewers screening and assessing bias with PROBAST+AI.

Analysis reported a difference of +13.0 percentage points for physicians minus AI (P = .038), and found image-only input was -25.9 percentage points lower than text-only (P = .001) and text-and-image was -10.6 points lower than text-only (P = .046).

What this doesn’t fix

Observed differences in accuracy by input modality may be confounded by task difficulty and information content, and evaluations lacked standardization.

Sources

  1. Peer-reviewedJapanese Journal of Radiology2026-10-03

The debate