TruaceTracing the truth around AISunday, September 20, 2026
Health·P Space·Evidence-backed problem·Published 2026-09-20

Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians

Abstract: Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs. Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Methods Thirty-nine frequently asked questions were submitted to each model. Responses w…

TRV-2026-1148Peer-reviewedPermanent record — cite & verify
Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians

Hospital Universitari Doctor Peset, València 10 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

The quick read

Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.

Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Responses were independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability; and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness.

Main points
  • Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information.
  • Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease).
  • Methods Thirty-nine frequently asked questions were submitted to each model.
Problem

Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.

The rundown

Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Methods Thirty-nine frequently asked questions were submitted to each model.

Sources

Reader signal

How should this claim be treated?

The debate