Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study
Purpose Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study. Methods In this controlled prompting study, 34 orthopaedic anae…
Guideline-augmented prompting increased blinded preference win rates and improved response consistency for LLMs answering orthopaedic anaesthesia questions compared to human experts.
LLM responses for orthopaedic anaesthesia showed variable consistency, with later models exhibiting greater partial variability despite avoiding fully contradictory outputs.
Comparative assessment relied on an LLM-as-a-judge rather than direct patient outcomes, and the authors note this method still needs validation.
Evidence
- Peer-reviewedJournal of Experimental Orthopaedics2026-10-03
How should this claim be treated?
Truvace Impact Record TRV-2026-1296, v1: “Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study.” Truvace, 2026-10-06. /record/TRV-2026-1296 (accessed at citation time). sha256 9c2d74501c2178c4…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1296 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace