AI-generated Reports in Medical Illustration (REMIL) from musculoskeletal radiology report text
Source article: "Reports in Medical Illustration (REMIL) in Musculoskeletal Radiology: An Evaluation of Evolving AI Models"
Abstract: Objective Radiology reports remain predominantly text-based, requiring clinicians and patients to mentally reconstruct imaging findings. Reports in Medical Illustration (REMIL) represent an emerging approach in which artificial intelligence (AI) generates simplified visual summaries directly from report text. This study aimed to evaluate the feasibility, anatomical accuracy, and clinical utility of AI-generated REMIL in musculoskeletal (MSK) radiology. Methods Twenty-five MSK imaging cases were selected. Identic…
Contested: both sides are scored from claims and sources, not community votes.
Hospital Universitari Doctor Peset, València 08 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
On September 4, 2026, a peer-reviewed study in Academic Radiology evaluated 25 musculoskeletal imaging cases where three AI systems generated visual summaries from report text alone. Two fellowship-trained MSK radiologists rated each image for anatomical accuracy and clinical usefulness, finding Gemini 3.0 Pro most consistent at 40-42% accurate and 60-65% useful, while ChatGPT and Perplexity frequently produced plausible but inaccurate images.
The findings matter because visual summaries could improve clinician and patient understanding of text-based radiology reports, but inconsistent accuracy creates risk if used without oversight. Uncertainty remains about performance beyond the small 25-case sample, across broader anatomies, and whether radiologist validation workflows can mitigate major errors in complex multi-structure cases.
- Study tested 25 MSK imaging cases using identical report text and standardized prompts across three premium multimodal AI systems.
- Two fellowship-trained musculoskeletal radiologists independently assessed each illustration for anatomical accuracy and clinical usefulness.
- Google Gemini 3.0 Pro outperformed ChatGPT (GPT-4 with DALL-E 3) and Perplexity AI, but still achieved anatomical accuracy in only 40-42% of cases.
AI systems can generate rapid visual summaries directly from musculoskeletal radiology report text, with the best model producing clinically useful images in a majority of tested cases.
Current AI models frequently produce visually plausible but anatomically inaccurate illustrations, with major errors across all models, making them unreliable for unsupervised clinical use.
The rundown
Researchers provided identical MSK report text and standardized prompts to ChatGPT with DALL-E 3, Perplexity AI, and Google Gemini 3.0 Pro to generate illustrations based solely on text, recording image-generation time and categorizing errors as minor or major.
Assessment by two fellowship-trained musculoskeletal radiologists found simpler cases with a single dominant abnormality were illustrated more accurately, while complex cases drove major errors, leading authors to conclude REMIL should be implemented only with radiologist validation.
Evaluation was limited to 25 selected MSK cases assessed by two radiologists, with performance dropping in complex cases involving multiple structures or planes.
Sources
- Peer-reviewedAcademic Radiology2026-09-04
How should this claim be treated?
ace
The debate