TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-06

AI-generated Reports in Medical Illustration (REMIL) from musculoskeletal radiology report text

Source article: "Reports in Medical Illustration (REMIL) in Musculoskeletal Radiology: An Evaluation of Evolving AI Models"

Abstract: Objective Radiology reports remain predominantly text-based, requiring clinicians and patients to mentally reconstruct imaging findings. Reports in Medical Illustration (REMIL) represent an emerging approach in which artificial intelligence (AI) generates simplified visual summaries directly from report text. This study aimed to evaluate the feasibility, anatomical accuracy, and clinical utility of AI-generated REMIL in musculoskeletal (MSK) radiology. Methods Twenty-five MSK imaging cases were selected. Identic…

TRV-2026-0996Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 71The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 67The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
"Reports in Medical Illustration (REMIL) in Musculoskeletal Radiology: An Evaluation of Evolving AI Models"

Hospital Universitari Doctor Peset, València 08 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

The quick read

On September 4, 2026, a peer-reviewed study in Academic Radiology evaluated 25 musculoskeletal imaging cases where three AI systems generated visual summaries from report text alone. Two fellowship-trained MSK radiologists rated each image for anatomical accuracy and clinical usefulness, finding Gemini 3.0 Pro most consistent at 40-42% accurate and 60-65% useful, while ChatGPT and Perplexity frequently produced plausible but inaccurate images.

The findings matter because visual summaries could improve clinician and patient understanding of text-based radiology reports, but inconsistent accuracy creates risk if used without oversight. Uncertainty remains about performance beyond the small 25-case sample, across broader anatomies, and whether radiologist validation workflows can mitigate major errors in complex multi-structure cases.

Main points
  • Study tested 25 MSK imaging cases using identical report text and standardized prompts across three premium multimodal AI systems.
  • Two fellowship-trained musculoskeletal radiologists independently assessed each illustration for anatomical accuracy and clinical usefulness.
  • Google Gemini 3.0 Pro outperformed ChatGPT (GPT-4 with DALL-E 3) and Perplexity AI, but still achieved anatomical accuracy in only 40-42% of cases.
Gain

AI systems can generate rapid visual summaries directly from musculoskeletal radiology report text, with the best model producing clinically useful images in a majority of tested cases.

Problem

Current AI models frequently produce visually plausible but anatomically inaccurate illustrations, with major errors across all models, making them unreliable for unsupervised clinical use.

The rundown

Researchers provided identical MSK report text and standardized prompts to ChatGPT with DALL-E 3, Perplexity AI, and Google Gemini 3.0 Pro to generate illustrations based solely on text, recording image-generation time and categorizing errors as minor or major.

Assessment by two fellowship-trained musculoskeletal radiologists found simpler cases with a single dominant abnormality were illustrated more accurately, while complex cases drove major errors, leading authors to conclude REMIL should be implemented only with radiologist validation.

What this doesn’t fix

Evaluation was limited to 25 selected MSK cases assessed by two radiologists, with performance dropping in complex cases involving multiple structures or planes.

Sources

Reader signal

How should this claim be treated?

The debate