LLM-based clinical note generation and summarisation and its impact on clinical documentation safety
Source article: A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation
Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM outputs, (2) an experimental structure for iterative comparisons in our LLM document generation pipeline, (3) a clinical sa…
Contested: both sides are scored from claims and sources, not community votes.

Hospital Doctor Peset. CC BY 3.0 · https://creativecommons.org/licenses/by/3.0
On 2025-05-13, a peer-reviewed framework was described for evaluating LLMs that automate summarising consultations into clinical notes. It combines an error taxonomy, iterative experimental comparisons, a clinical safety harm assessment, and the CREOLA interface, tested across 18 configurations with 12,999 clinician-annotated sentences.
The study matters because it moves beyond anecdotal accuracy claims to measured safety rates, reporting 1.47% hallucinations and 3.45% omissions while demonstrating that targeted refinements can lower major errors below human benchmarks. Uncertainty remains about performance outside the tested configurations, note types, and clinical settings.
- Proposes framework with error taxonomy, experimental structure for iterative comparisons, clinical safety framework, and CREOLA graphical user interface.
- Evaluation derived from 18 experimental configurations for clinical note generation comprising 12,999 clinician-annotated sentences.
- Observed 1.47% hallucination rate and 3.45% omission rate, with iterative refinement reducing major errors below reported human rates.
Refining prompts and workflows within the proposed framework reduced major errors below previously reported human note-taking rates, supporting safer clinical documentation.
LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.
The rundown
The authors built a four-part framework to assess LLM outputs for summarising consultations: an error taxonomy, an experimental structure for iterative comparisons in their document generation pipeline, a clinical safety framework to evaluate harms, and a GUI called CREOLA to facilitate the process. Clinical error metrics were calculated from 18 configurations producing clinical notes, totaling 12,999 sentences annotated by clinicians.
Results quantified fidelity gaps at 1.47% hallucination and 3.45% omission, and showed that prompt and workflow refinements could drive major errors below previously reported human note-taking rates. The work positions the framework as a method to iteratively measure and improve safety of automated clinical documentation.
Findings are bounded to 18 experimental configurations and 12,999 clinician-annotated sentences for clinical note generation, limiting generalizability beyond that evaluation set.
Sources
- Peer-reviewednpj Digital Medicine2025-05-13
How should this claim be treated?
ace
The debate