TruaceTracing the truth around AISaturday, September 12, 2026
TRV-2026-1018Certified recordPeer-reviewed

Select large language models outperform hip preservation experts on consensus-based hip preservation questionnaire

Artificial intelligence (AI) is increasingly utilized in medical education and clinical contexts, yet few studies compare the performance of large language models (LLMs) to subspecialized experts in providing guideline-based medical information on hip preservation. The purpose of this study was to evaluate the performance of three LLMs compared to a panel of international hip preservation experts in answering guideline-based questions related to femoroacetabular impingement syndrome, hip dysplasia and microinsta…

Health · The Trace — both readings · certified 2026-09-08 · v1 · article view · machine-readable

Current reading — gain

Three large language models achieved higher accuracy than a panel of hip preservation experts on a 21-item consensus-based questionnaire covering femoroacetabular impingement syndrome, hip dysplasia and microinstability.

Current reading — problem

Even when incorrect, ChatGPT and Claude produced thorough justifications, creating risk of convincing but wrong guideline-based information, while Gemini showed formatting deviations.

What this doesn’t fix

Findings are based on a small expert sample of 10 respondents and a structured 21-item verifiable question set, limiting generalizability to real-world clinical decision-making, and the study is labeled Level V evidence.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-1018, v1: “Select large language models outperform hip preservation experts on consensus-based hip preservation questionnaire.” Truvace, 2026-09-08. /record/TRV-2026-1018 (accessed at citation time). sha256 a9faadcee3050f7f

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv1a9faadcee305

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1018 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.