TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-10

safety and quality of AI-generated patient-friendly radiology reports for patient distribution

Source article: A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports

Abstract: Background: Artificial intelligence (AI) tools are being used to translate radiology reports into plain language, but translation errors may compromise comprehension and safety. Objective: To develop and evaluate a rubric for assessing the quality and safety of AI-generated patient-friendly radiology reports. Methods: In this prospective study (conducted from February 2025 to December 2025), survey-workshop cycles, involving lay participants and a multidisciplinary panel, were used to develop a rubric for gradin…

TRV-2026-1041Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 73The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 69The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports

United States Naval Medical Bulletin Vol. 44, Nos. 1-6, 1945 by U.S. Navy. Bureau of Medicine and Surgery. Public domain

The quick read

From February to December 2025, researchers developed a 5-attribute rubric  clarity, content, certainty, tone, verbosity  to grade AI-generated patient-friendly radiology reports, using survey-workshop cycles with 19 participants and testing with ChatGPT-4.1 and Claude-4.0 outputs from public radiology impressions. Evaluation involved six research-team members and 111 additional participants, plus AI evaluation with ChatGPT-5, comparing rubric grades to prespecified reference standards and to subjective decisions about withholding unsafe reports.

The work matters because health systems are already using AI to simplify radiology reports for patients, where errors can affect understanding and safety, and a standardized check could enable scalable quality assurance. Uncertainty remains because agreement with reference standards was only moderate in wider field testing, AI grading reached only moderate agreement, and authors state further training and validation are needed before clinical integration.

Main points
  • Prospective study from February 2025 to December 2025 used survey-workshop cycles with lay participants and multidisciplinary panel to develop rubric.
  • Final rubric grades five core attributes  clarity, content, certainty, tone, verbosity  on 3-point scale; grade 1 in any attribute other than verbosity means unsafe for distribution.
  • Lay and radiologist research-team members had almost-perfect intergroup agreement =0.87 for overall grades across 60 reports.
  • In field testing, 80 lay participants evaluating 480 reports had moderate agreement =0.43 with reference-standard grades; AI evaluation had =0.44 and 88.1% agreement on rule-based distribution decisions.
Gain

A five-attribute rubric for AI-generated patient-friendly radiology reports showed almost-perfect agreement between lay and radiologist team members and may provide a standardized safeguard before patient distribution.

Problem

AI tools translating radiology reports into plain language can produce translation errors that compromise comprehension and safety, causing reports to be graded unsafe and warrant withholding from patients.

The rundown

Researchers generated patient-friendly versions of radiology impressions from a public dataset using ChatGPT-4.1 and Claude-4.0 with prespecified quality targets, then had research-team members, additional lay and radiologist participants, and ChatGPT-5 evaluate them using the rubric.

Additional lay (n=19) and radiologist (n=12) participants each evaluating six reports showed 91.2% and 95.8% agreement between subjective and rubric rule-based distribution decisions, while wider testing with 80 lay participants showed 73.5% agreement between subjective and rule-based decisions.

What this doesn’t fix

Authors note rubric requires further training and validation before clinical use, and wider field testing showed only moderate agreement with reference standards.

Sources

Reader signal

How should this claim be treated?

The debate