TruaceTracing the truth around AISaturday, September 12, 2026
TRV-2026-0996Version 1 · Certified

Written 2026-09-06 06:06:37 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0996
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-06T06:06:37.125715Z
status: published
lens: trace
sector: health
headline: "Reports in Medical Illustration (REMIL) in Musculoskeletal Radiology: An Evaluation of Evolving AI Models"
dek: Objective Radiology reports remain predominantly text-based, requiring clinicians and patients to mentally reconstruct imaging findings. Reports in Medical Illustration (REMIL) represent an emerging approach in which artificial intelligence (AI) generates simplified visual summaries directly from report text. This study aimed to evaluate the feasibility, anatomical accuracy, and clinical utility of AI-generated REMIL in musculoskeletal (MSK) radiology. Methods Twenty-five MSK imaging cases were selected. Identic…
gain_title: AI systems can generate rapid visual summaries directly from musculoskeletal radiology report text, with the best model producing clinically useful images in a majority of tested cases.
problem_title: Current AI models frequently produce visually plausible but anatomically inaccurate illustrations, with major errors across all models, making them unreliable for unsupervised clinical use.
trace_subject: AI-generated Reports in Medical Illustration (REMIL) from musculoskeletal radiology report text
gain_reading: AI systems can generate rapid visual summaries directly from musculoskeletal radiology report text, with the best model producing clinically useful images in a majority of tested cases.
gain_evidence: Google Gemini 3.0 Pro demonstrated the most consistent performance, producing anatomically accurate illustrations in approximately 40-42% of cases and clinically useful images in 60-65% of cases
problem_reading: Current AI models frequently produce visually plausible but anatomically inaccurate illustrations, with major errors across all models, making them unreliable for unsupervised clinical use.
problem_evidence: current AI models exhibit inconsistent anatomical accuracy and are not yet reliable for unsupervised clinical use | Major errors were observed across all models, particularly in complex cases involving multiple anatomical structures or imaging planes
quick_read: On September 4, 2026, a peer-reviewed study in Academic Radiology evaluated 25 musculoskeletal imaging cases where three AI systems generated visual summaries from report text alone. Two fellowship-trained MSK radiologists rated each image for anatomical accuracy and clinical usefulness, finding Gemini 3.0 Pro most consistent at 40-42% accurate and 60-65% useful, while ChatGPT and Perplexity frequently produced plausible but inaccurate images.

The findings matter because visual summaries could improve clinician and patient understanding of text-based radiology reports, but inconsistent accuracy creates risk if used without oversight. Uncertainty remains about performance beyond the small 25-case sample, across broader anatomies, and whether radiologist validation workflows can mitigate major errors in complex multi-structure cases.
limitation: Evaluation was limited to 25 selected MSK cases assessed by two radiologists, with performance dropping in complex cases involving multiple structures or planes.
tag: Dual reading
key_points: Study tested 25 MSK imaging cases using identical report text and standardized prompts across three premium multimodal AI systems. | Two fellowship-trained musculoskeletal radiologists independently assessed each illustration for anatomical accuracy and clinical usefulness. | Google Gemini 3.0 Pro outperformed ChatGPT (GPT-4 with DALL-E 3) and Perplexity AI, but still achieved anatomical accuracy in only 40-42% of cases.
rundown: Researchers provided identical MSK report text and standardized prompts to ChatGPT with DALL-E 3, Perplexity AI, and Google Gemini 3.0 Pro to generate illustrations based solely on text, recording image-generation time and categorizing errors as minor or major.

Assessment by two fellowship-trained musculoskeletal radiologists found simpler cases with a single dominant abnormality were illustrated more accurately, while complex cases drove major errors, leading authors to conclude REMIL should be implemented only with radiologist validation.
sources:
- peer_reviewed | Academic Radiology | https://doi.org/10.1016/j.acra.2026.08.082 | 2026-09-04
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
ce033e6381c09b311e47ffeaf84d1b6d1e91e03fdc881409132fc74f376fc028
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0996 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.