TruaceTracing the truth around AISaturday, September 12, 2026
TRV-2026-1041Version 1 · Certified

Written 2026-09-10 06:04:33 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1041
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-10T06:04:33.355750Z
status: published
lens: trace
sector: health
headline: A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports
dek: Background: Artificial intelligence (AI) tools are being used to translate radiology reports into plain language, but translation errors may compromise comprehension and safety. Objective: To develop and evaluate a rubric for assessing the quality and safety of AI-generated patient-friendly radiology reports. Methods: In this prospective study (conducted from February 2025 to December 2025), survey-workshop cycles, involving lay participants and a multidisciplinary panel, were used to develop a rubric for gradin…
gain_title: A five-attribute rubric for AI-generated patient-friendly radiology reports showed almost-perfect agreement between lay and radiologist team members and may provide a standardized safeguard before patient distribution.
problem_title: AI tools translating radiology reports into plain language can produce translation errors that compromise comprehension and safety, causing reports to be graded unsafe and warrant withholding from patients.
trace_subject: safety and quality of AI-generated patient-friendly radiology reports for patient distribution
gain_reading: A five-attribute rubric for AI-generated patient-friendly radiology reports showed almost-perfect agreement between lay and radiologist team members and may provide a standardized safeguard before patient distribution.
gain_evidence: The rubric may provide a standardized safeguard before release of AI-generated patient-friendly reports. | AI rubric application could enable scalable quality assurance and safer clinical integration of AI-generated communications.
problem_reading: AI tools translating radiology reports into plain language can produce translation errors that compromise comprehension and safety, causing reports to be graded unsafe and warrant withholding from patients.
problem_evidence: translation errors may compromise comprehension and safety | patient-friendly reports assessed as grade 1 (unsafe or unacceptable) in any attribute other than verbosity are unsafe for distribution and warrant withholding
quick_read: From February to December 2025, researchers developed a 5-attribute rubric  clarity, content, certainty, tone, verbosity  to grade AI-generated patient-friendly radiology reports, using survey-workshop cycles with 19 participants and testing with ChatGPT-4.1 and Claude-4.0 outputs from public radiology impressions. Evaluation involved six research-team members and 111 additional participants, plus AI evaluation with ChatGPT-5, comparing rubric grades to prespecified reference standards and to subjective decisions about withholding unsafe reports.

The work matters because health systems are already using AI to simplify radiology reports for patients, where errors can affect understanding and safety, and a standardized check could enable scalable quality assurance. Uncertainty remains because agreement with reference standards was only moderate in wider field testing, AI grading reached only moderate agreement, and authors state further training and validation are needed before clinical integration.
limitation: Authors note rubric requires further training and validation before clinical use, and wider field testing showed only moderate agreement with reference standards.
tag: Dual reading
key_points: Prospective study from February 2025 to December 2025 used survey-workshop cycles with lay participants and multidisciplinary panel to develop rubric. | Final rubric grades five core attributes  clarity, content, certainty, tone, verbosity  on 3-point scale; grade 1 in any attribute other than verbosity means unsafe for distribution. | Lay and radiologist research-team members had almost-perfect intergroup agreement =0.87 for overall grades across 60 reports. | In field testing, 80 lay participants evaluating 480 reports had moderate agreement =0.43 with reference-standard grades; AI evaluation had =0.44 and 88.1% agreement on rule-based distribution decisions.
rundown: Researchers generated patient-friendly versions of radiology impressions from a public dataset using ChatGPT-4.1 and Claude-4.0 with prespecified quality targets, then had research-team members, additional lay and radiologist participants, and ChatGPT-5 evaluate them using the rubric.

Additional lay (n=19) and radiologist (n=12) participants each evaluating six reports showed 91.2% and 95.8% agreement between subjective and rubric rule-based distribution decisions, while wider testing with 80 lay participants showed 73.5% agreement between subjective and rule-based decisions.
sources:
- peer_reviewed | American Journal of Roentgenology | https://doi.org/10.2214/ajr.26.35532 | 2026-09-09
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
63a875426e2aeed18536ae8d73b3a85a71e2782c6d7fd1e0df8992f318015159
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1041 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.