TruaceTracing the truth around AIMonday, September 14, 2026
TRV-2026-1069Version 1 · Certified

Written 2026-09-13 06:56:20 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1069
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-13T06:56:20.753741Z
status: published
lens: trace
sector: health
headline: Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty
dek: Purpose To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). Methods Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments,…
gain_title: Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.
problem_title: LLM-generated answers to patient questions about robotic-assisted total knee arthroplasty remained above recommended patient-education reading levels and should be regarded as supplementary rather than standalone sources of information.
trace_subject: LLM-generated patient information about robotic-assisted total knee arthroplasty
gain_reading: Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.
gain_evidence: The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. | Median DISCERN scores were 46.0 (range, 35.0-50.0) for ChatGPT-o3, 45.75 (28.5-51.0) for ChatGPT-5.2, 43.75 (32.0-47.5) for Gemini 3, and 42.0 (32.0-50.0) for DeepSeek
problem_reading: LLM-generated answers to patient questions about robotic-assisted total knee arthroplasty remained above recommended patient-education reading levels and should be regarded as supplementary rather than standalone sources of information.
problem_evidence: median responses across all models remained above recommended patient-education reading levels. | LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.
quick_read: A September 2026 peer-reviewed study in The Knee compared four large language models on 30 common patient questions about robotic-assisted total knee arthroplasty, evaluating responses with DISCERN, QAMAI, a 5-point clinical accuracy scale, and PEMAT and Flesch-Kincaid readability measures.

The results matter because patients increasingly use LLMs for surgical information, yet the study found generally acceptable accuracy paired with readability above recommended levels and variable quality scores, leaving uncertainty about how to integrate these tools safely into preoperative education without replacing clinician counseling.
limitation: Findings are limited to 30 selected questions and show readability remains above recommended patient-education levels, with no significant difference on QAMAI and authors concluding responses should be supplementary only.
tag: Dual reading
key_points: Study compared ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek on 30 patient questions about RA-TKA identified via LLM outputs and Google search queries. | Evaluation used DISCERN and QAMAI for information quality, a 5-point ordinal scale for clinical accuracy, and PEMAT Understandability and Flesch-Kincaid Reading Ease for readability. | Median DISCERN scores ranged from 42.0 to 46.0 across models with a significant overall difference among models. | No significant difference was detected using QAMAI, while Gemini 3 and DeepSeek showed greater readability in selected comparisons.
rundown: Researchers identified 30 frequently asked patient questions about robotic-assisted total knee arthroplasty using LLM outputs and Google search queries, then generated responses from ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek.

Responses were scored with DISCERN and QAMAI for quality, a 5-point ordinal rating for clinical accuracy, and PEMAT Understandability and Flesch-Kincaid Reading Ease for readability, finding median DISCERN scores from 42.0 to 46.0 with significant differences among models.

The authors concluded that while clinical accuracy was generally acceptable and Gemini 3 and DeepSeek were more readable in some comparisons, all models exceeded recommended reading levels and should be considered supplementary rather than standalone patient education.
sources:
- peer_reviewed | The Knee | https://doi.org/10.1016/j.knee.2026.104635 | 2026-09-11
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
9881c86c0a3df0bca67dfe02cce84837973515c8ea8df8bd3479582447d3d8d2
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1069 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.