TruaceTracing the truth around AIWednesday, August 5, 2026
TRV-2026-0522Certified recordPeer-reviewed

A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation

Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM outputs, (2) an experimental structure for iterative comparisons in our LLM document generation pipeline, (3) a clinical sa…

Health · The Trace — both readings · certified 2026-07-24 · v1 · article view · machine-readable

Current reading — gain

Refining prompts and workflows within the proposed framework reduced major errors below previously reported human note-taking rates, supporting safer clinical documentation.

Current reading — problem

LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.

What this doesn’t fix

Findings are bounded to 18 experimental configurations and 12,999 clinician-annotated sentences for clinical note generation, limiting generalizability beyond that evaluation set.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-0522, v1: “A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation.” Truvace, 2026-07-24. /record/TRV-2026-0522 (accessed at citation time). sha256 5ce93b85276c3786

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv15ce93b85276c

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0522 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.