TruaceTracing the truth around AISunday, September 20, 2026
TRV-2026-1148Version 1 · Certified

Written 2026-09-20 06:53:25 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1148
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-20T06:53:25.808142Z
status: published
lens: p_space
sector: health
headline: Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians
dek: Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs. Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Methods Thirty-nine frequently asked questions were submitted to each model. Responses w…
gain_title: (none)
problem_title: Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.
trace_subject: (none)
gain_reading: (none)
gain_evidence: (none)
problem_reading: Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.
problem_evidence: (none)
quick_read: Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.

Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Responses were independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability; and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness.
limitation: 
tag: Evidence-backed problem
key_points: Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. | Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). | Methods Thirty-nine frequently asked questions were submitted to each model.
rundown: Background Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.

Aim To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease). Methods Thirty-nine frequently asked questions were submitted to each model.
sources:
- peer_reviewed | World Journal of Methodology | https://doi.org/10.5662/wjm.v16.i3.116022 | 2026-09-20
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
3d673ec1409fda44c919c155b6b47a2eab46c633bf04f331ea4672ff395984bd
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1148 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.