automated identification of cardiovascular events from EMRs using LLM-assisted workflow versus ICD-based retrieval
Source article: Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States
Abstract: To compare the diagnostic accuracy of four available automated electronic medical record (EMR) retrieval methods, including a large language model (LLM)-assisted workflow, against manual chart adjudication for identifying cardiovascular events. Retrospective diagnostic accuracy study. Three sites within a single US tertiary health system. Two adult cohorts with previously adjudicated cardiovascular outcomes were included. Cohort 1 included 2258 patients treated with immune checkpoint inhibitors, and Cohort 2 inc…
Contested: both sides are scored from claims and sources, not community votes.
Hospital Universitari Doctor Peset, València 09 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
A multisite retrospective validation study in a US tertiary health system compared four automated EMR retrieval methods to manual chart adjudication for ischaemic stroke/TIA, MI, HF exacerbation/hospitalisation and composite MACE in 2258 patients treated with immune checkpoint inhibitors and 1426 patients who underwent TAVR. The zero-shot LLM workflow achieved the highest AUCs for most outcomes, while ICD-based retrieval remained competitive.
The findings matter because accurate automated capture of cardiovascular events is critical for retrospective outcomes research and pharmacovigilance, and the results suggest LLM extraction can complement but not uniformly replace structured code-based methods. Uncertainty remains about generalizability beyond a single health system, performance across other populations and outcomes, and prospective implementation costs and workflow integration.
- Retrospective diagnostic accuracy study compared four automated EMR retrieval methods against clinician manual chart adjudication as reference standard.
- Two adult cohorts from three sites within a single US tertiary health system: 2258 patients treated with immune checkpoint inhibitors and 1426 patients who underwent transcatheter aortic valve replacement.
- Outcomes evaluated were ischaemic stroke or transient ischaemic attack, myocardial infarction, heart failure exacerbation or hospitalisation, and composite MACE.
- AUC, sensitivity, specificity and net reclassification improvement were assessed for ICD codes, primary diagnosis, problem list and zero-shot LLM workflow.
A zero-shot LLM-assisted workflow improved automated identification of cardiovascular events from EMRs, achieving the highest AUCs for stroke, MI and composite MACE in two cohorts compared to ICD codes, primary diagnosis and problem list methods.
The LLM-assisted workflow did not consistently outperform ICD-based retrieval, with no statistically significant AUC difference in Cohort 1 and lower AUC than ICD codes for heart failure in that cohort.
The rundown
Researchers compared ICD codes, primary diagnosis, problem list, and a zero-shot LLM workflow against clinician manual chart adjudication in 2258 immune checkpoint inhibitor patients and 1426 TAVR patients across three sites.
LLM had highest AUCs for most outcomes, including 0.920 for stroke and 0.938 for MI in Cohort 1 and 0.915 for stroke and 0.928 for MI in Cohort 2, but ICD codes had higher AUC for HF in Cohort 1 and differences were not statistically significant in that cohort, while LLM was significantly higher for stroke and composite MACE in Cohort 2.
Retrospective validation limited to three sites within a single US tertiary health system with two specific cohorts, and performance was context-dependent with ICD-based retrieval remaining competitive for some outcomes.
Sources
- Peer-reviewedBMJ Open2026-08-13
How should this claim be treated?
ace
The debate