TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-08

accuracy and reliability of LLMs versus hip preservation experts when answering consensus-based hip preservation questions

Source article: Select large language models outperform hip preservation experts on consensus-based hip preservation questionnaire

Abstract: Artificial intelligence (AI) is increasingly utilized in medical education and clinical contexts, yet few studies compare the performance of large language models (LLMs) to subspecialized experts in providing guideline-based medical information on hip preservation. The purpose of this study was to evaluate the performance of three LLMs compared to a panel of international hip preservation experts in answering guideline-based questions related to femoroacetabular impingement syndrome, hip dysplasia and microinsta…

TRV-2026-1018Peer-reviewedPermanent record — cite & verify
Trace impact reading

Negative state: both sides are scored from claims and sources, not community votes.

P 71The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 66The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Select large language models outperform hip preservation experts on consensus-based hip preservation questionnaire

Physical-Therapy-Dept,-Deshon-General-Hospital,-West-end-whirlpool-room by The US government, most likely the US National Institutes of Health .. Public domain

The quick read

A peer-reviewed study published September 7, 2026 compared three large language models to 10 international hip preservation experts on a 21-item questionnaire derived from consensus guidelines for femoroacetabular impingement, dysplasia and microinstability. Gemini scored 100%, ChatGPT 98.4% and Claude 96.8% versus 90.5% for experts, with models also showing higher intra-item agreement.

The result suggests newer LLMs can reliably reproduce structured, verifiable orthopaedic consensus knowledge and may support education, but the observed pattern of thorough justifications for incorrect answers and formatting deviations raises concerns about overtrust in clinical contexts. Uncertainty remains because the test was limited to a small expert panel and a controlled questionnaire rather than live patient care.

Main points
  • Study compared ChatGPT 5.2, Gemini 3 and Claude 4.5 Sonnet to 10 international hip preservation specialists on a 21-item questionnaire based on published consensus guidelines.
  • Expert accuracy was 90.5% with 42.9% percent agreement and Fleiss' ba 0.769, while AI intra-item agreement reached 100% for Gemini and 95.2% for ChatGPT and Claude.
  • Authors concluded LLMs are likely to serve as an adjunct in orthopaedic education and practice but noted need for rigorous exploration of limitations.
Gain

Three large language models achieved higher accuracy than a panel of hip preservation experts on a 21-item consensus-based questionnaire covering femoroacetabular impingement syndrome, hip dysplasia and microinstability.

Problem

Even when incorrect, ChatGPT and Claude produced thorough justifications, creating risk of convincing but wrong guideline-based information, while Gemini showed formatting deviations.

The rundown

Researchers built a 21-item questionnaire from published consensus guidelines on femoroacetabular impingement syndrome, hip dysplasia and microinstability, then surveyed 10 hip preservation specialists and prompted three LLMs with three runs per item to measure accuracy, percent agreement, Fleiss' ba and justification quality.

Results by the September 2026 publication showed Gemini at 100%, ChatGPT at 98.4% and Claude at 96.8% accuracy versus 90.5% for experts, with higher intra-item consistency for models, but qualitative review flagged convincing justifications for wrong answers and formatting issues.

What this doesn’t fix

Findings are based on a small expert sample of 10 respondents and a structured 21-item verifiable question set, limiting generalizability to real-world clinical decision-making, and the study is labeled Level V evidence.

Reader signal

How should this claim be treated?

The debate