TruaceTracing the truth around AITuesday, July 21, 2026
Health·The Trace·Automated dual reading·Published 2026-07-20

accuracy of ChatGPT answers to common patient hip arthroscopy questions

Source article: ChatGPT Provides Satisfactory but Occasionally Inaccurate Answers to Common Patient Hip Arthroscopy Questions

PURPOSE: To assess the ability of ChatGPT to answer common patient questions regarding hip arthroscopy, and to analyze the accuracy and appropriateness of its responses. METHODS: Ten questions were selected from well-known patient education websites, and ChatGPT (version 3.5) responses to these questions were graded by 2 fellowship-trained hip preservation surgeons. Responses were analyzed, compared with the current literature, and graded from A to D (A being the highest, and D being the lowest) in a grading sca…

TRV-2026-0412Peer-reviewedPermanent record — cite & verify
Trace impact reading

Negative state: both sides are scored from claims and sources, not community votes.

P 73The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 68The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
ChatGPT Provides Satisfactory but Occasionally Inaccurate Answers to Common Patient Hip Arthroscopy Questions

"student_ipad_school - 115" by flickingerbrad is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.

The quick read

In a study published June 22, 2024, two hip preservation surgeons graded ChatGPT 3.5 answers to ten common hip arthroscopy questions drawn from patient education sites, using an A-to-D scale and readability scores FRES and FKGL.

The findings matter because patients increasingly turn to chatbots for surgical information, yet the mix of mostly satisfactory grades alongside documented inaccuracies and college-level readability raises questions about safe use without clinician oversight.

Main points
  • Ten common patient questions from education websites were answered by ChatGPT version 3.5 and graded A to D by two fellowship-trained hip preservation surgeons.
  • Consensus grades were A 50%, B 30%, C 10%, D 10%, with initial inter-rater agreement of 30%.
  • Readability analysis found mean Flesch-Kincaid Reading Ease Score 28.2 and mean Grade Level 14.4, indicating college-level text.
Gain

ChatGPT provided satisfactory answers to common patient hip arthroscopy questions, with half of responses graded A and another 30% graded B by fellowship-trained surgeons.

Problem

ChatGPT responses contained incorrect information in more than one instance and were written at a college graduate reading level, requiring caution for patient education.

The rundown

Researchers selected ten common hip arthroscopy questions from patient education websites and prompted ChatGPT 3.5, then had two fellowship-trained hip preservation surgeons grade answers A to D for accuracy and completeness, reaching consensus when needed.

Results showed five As, three Bs, one C and one D, with mean Flesch-Kincaid Reading Ease 28.2 and Grade Level 14.4, and authors noted potential to aid physicians while warning about inaccuracies.

Sources

Reader signal

How should this claim be treated?

The debate