Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians
Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs. Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Methods Thirty-nine frequently asked questions were submitted to each model. Responses w…
Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.
Evidence
- Peer-reviewedWorld Journal of Methodology2026-09-20
How should this claim be treated?
Truvace Impact Record TRV-2026-1148, v1: “Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians.” Truvace, 2026-09-20. /record/TRV-2026-1148 (accessed at citation time). sha256 3d673ec1409fda44…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1148 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace