LLM-generated answers to orthopaedic anaesthesia questions evaluated for preference and consistency

Source article: Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study

Purpose Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study. Methods In this controlled prompting study, 34 orthopaedic anae…

Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study
Physical therapy in a pool agitator (SC 495871), National Museum of Health and Medicine (3300120120) by National Museum of Health and Medicine. CC BY 2.0 · https://creativecommons.org/licenses/by/2.0
Trace impact readingContested
P 69The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.

Both sides are scored from claims and sources, not community votes.

G 71The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.

In brief

In a controlled prompting study of 34 orthopaedic anaesthesia questions, three LLMs and three expert anesthesiologists generated answers, with GPT-4o and GPT-5.2 tested both with and without clinical practice guidelines. An independent guideline-informed LLM judge performed blinded pairwise comparisons, recording preference, consistency and confidence.

The findings matter because orthopaedic anaesthesia involves high comorbidity burden and guideline-dependent management where LLM assistance is being considered. While successive generations and guideline augmentation improved preference and consistency, variability persisted and reliance on an LLM judge leaves uncertainty about real-world clinical safety and generalizability.

Main points

  1. Controlled prompting study used 34 orthopaedic anaesthesia questions spanning seven clinical subdomains answered by GPT-3.5-turbo, GPT-4o, GPT-5.2 and three expert anesthesiologists.
  2. Blinded pairwise evaluation was performed by an independent guideline-informed LLM-as-a-judge (GPT-5.2) using same guideline material as reference standard.
  3. GPT-4o showed the highest baseline consistency at 80.4% while GPT-5.2 showed no fully contradictory outputs but greater partial variability before augmentation.

The gain

Guideline-augmented prompting increased blinded preference win rates and improved response consistency for LLMs answering orthopaedic anaesthesia questions compared to human experts.

The problem

LLM responses for orthopaedic anaesthesia showed variable consistency, with later models exhibiting greater partial variability despite avoiding fully contradictory outputs.

The rundown

The study tested three successive generations with and without access to relevant clinical practice guidelines, anonymizing all outputs for blinded comparison against human expert answers.

Intra-model consistency was measured from repeated independent generations, and qualitative analysis linked preference for later generations to greater completeness, explicit clinical reasoning and guideline-oriented responses.

What this doesn’t fix

Comparative assessment relied on an LLM-as-a-judge rather than direct patient outcomes, and the authors note this method still needs validation.

Sources

  1. Peer-reviewedJournal of Experimental Orthopaedics2026-10-03

The debate