LLM-generated patient information about robotic-assisted total knee arthroplasty
Source article: Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty
Abstract: Purpose To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). Methods Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments,…
Contested: both sides are scored from claims and sources, not community votes.

"student_ipad_school - 078" by flickingerbrad is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
A September 2026 peer-reviewed study in The Knee compared four large language models on 30 common patient questions about robotic-assisted total knee arthroplasty, evaluating responses with DISCERN, QAMAI, a 5-point clinical accuracy scale, and PEMAT and Flesch-Kincaid readability measures.
The results matter because patients increasingly use LLMs for surgical information, yet the study found generally acceptable accuracy paired with readability above recommended levels and variable quality scores, leaving uncertainty about how to integrate these tools safely into preoperative education without replacing clinician counseling.
- Study compared ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek on 30 patient questions about RA-TKA identified via LLM outputs and Google search queries.
- Evaluation used DISCERN and QAMAI for information quality, a 5-point ordinal scale for clinical accuracy, and PEMAT Understandability and Flesch-Kincaid Reading Ease for readability.
- Median DISCERN scores ranged from 42.0 to 46.0 across models with a significant overall difference among models.
- No significant difference was detected using QAMAI, while Gemini 3 and DeepSeek showed greater readability in selected comparisons.
Large language models including ChatGPT-o3, ChatGPT-5.2, Gemini 3 and DeepSeek provided generally acceptable clinical accuracy when answering 30 frequently asked patient questions about robotic-assisted total knee arthroplasty.
LLM-generated answers to patient questions about robotic-assisted total knee arthroplasty remained above recommended patient-education reading levels and should be regarded as supplementary rather than standalone sources of information.
The rundown
Researchers identified 30 frequently asked patient questions about robotic-assisted total knee arthroplasty using LLM outputs and Google search queries, then generated responses from ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek.
Responses were scored with DISCERN and QAMAI for quality, a 5-point ordinal rating for clinical accuracy, and PEMAT Understandability and Flesch-Kincaid Reading Ease for readability, finding median DISCERN scores from 42.0 to 46.0 with significant differences among models.
The authors concluded that while clinical accuracy was generally acceptable and Gemini 3 and DeepSeek were more readable in some comparisons, all models exceeded recommended reading levels and should be considered supplementary rather than standalone patient education.
Findings are limited to 30 selected questions and show readability remains above recommended patient-education levels, with no significant difference on QAMAI and authors concluding responses should be supplementary only.
Sources
- Peer-reviewedThe Knee2026-09-11
How should this claim be treated?
ace
The debate