TruaceTracing the truth around AIWednesday, August 5, 2026
TRV-2026-0522Version 1 · Certified

Written 2026-07-24 00:24:28 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0522
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-24T00:24:28.692134Z
status: published
lens: trace
sector: health
headline: A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation
dek: Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM outputs, (2) an experimental structure for iterative comparisons in our LLM document generation pipeline, (3) a clinical sa…
gain_title: Refining prompts and workflows within the proposed framework reduced major errors below previously reported human note-taking rates, supporting safer clinical documentation.
problem_title: LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.
trace_subject: LLM-based clinical note generation and summarisation and its impact on clinical documentation safety
gain_reading: Refining prompts and workflows within the proposed framework reduced major errors below previously reported human note-taking rates, supporting safer clinical documentation.
gain_evidence: "successfully reduced major errors below previously reported human note-taking rates" | "potential for safer clinical documentation"
problem_reading: LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.
problem_evidence: "1.47% hallucination rate" | "3.45% omission rate" | "compromise in patient safety"
quick_read: On 2025-05-13, a peer-reviewed framework was described for evaluating LLMs that automate summarising consultations into clinical notes. It combines an error taxonomy, iterative experimental comparisons, a clinical safety harm assessment, and the CREOLA interface, tested across 18 configurations with 12,999 clinician-annotated sentences.

The study matters because it moves beyond anecdotal accuracy claims to measured safety rates, reporting 1.47% hallucinations and 3.45% omissions while demonstrating that targeted refinements can lower major errors below human benchmarks. Uncertainty remains about performance outside the tested configurations, note types, and clinical settings.
limitation: Findings are bounded to 18 experimental configurations and 12,999 clinician-annotated sentences for clinical note generation, limiting generalizability beyond that evaluation set.
tag: Automated dual reading
key_points: Proposes framework with error taxonomy, experimental structure for iterative comparisons, clinical safety framework, and CREOLA graphical user interface. | Evaluation derived from 18 experimental configurations for clinical note generation comprising 12,999 clinician-annotated sentences. | Observed 1.47% hallucination rate and 3.45% omission rate, with iterative refinement reducing major errors below reported human rates.
rundown: The authors built a four-part framework to assess LLM outputs for summarising consultations: an error taxonomy, an experimental structure for iterative comparisons in their document generation pipeline, a clinical safety framework to evaluate harms, and a GUI called CREOLA to facilitate the process. Clinical error metrics were calculated from 18 configurations producing clinical notes, totaling 12,999 sentences annotated by clinicians.

Results quantified fidelity gaps at 1.47% hallucination and 3.45% omission, and showed that prompt and workflow refinements could drive major errors below previously reported human note-taking rates. The work positions the framework as a method to iteratively measure and improve safety of automated clinical documentation.
sources:
- peer_reviewed | npj Digital Medicine | https://doi.org/10.1038/s41746-025-01670-7 | 2025-05-13
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
5ce93b85276c378616d484bead12409019a3eee3c096d56db4ef612093bdcbd9
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0522 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.