TruaceTracing the truth around AITuesday, August 25, 2026
TRV-2026-0790Version 1 · Certified

Written 2026-08-16 06:23:20 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0790
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-08-16T06:23:20.229877Z
status: published
lens: trace
sector: health
headline: Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States
dek: To compare the diagnostic accuracy of four available automated electronic medical record (EMR) retrieval methods, including a large language model (LLM)-assisted workflow, against manual chart adjudication for identifying cardiovascular events. Retrospective diagnostic accuracy study. Three sites within a single US tertiary health system. Two adult cohorts with previously adjudicated cardiovascular outcomes were included. Cohort 1 included 2258 patients treated with immune checkpoint inhibitors, and Cohort 2 inc…
gain_title: A zero-shot LLM-assisted workflow improved automated identification of cardiovascular events from EMRs, achieving the highest AUCs for stroke, MI and composite MACE in two cohorts compared to ICD codes, primary diagnosis and problem list methods.
problem_title: The LLM-assisted workflow did not consistently outperform ICD-based retrieval, with no statistically significant AUC difference in Cohort 1 and lower AUC than ICD codes for heart failure in that cohort.
trace_subject: automated identification of cardiovascular events from EMRs using LLM-assisted workflow versus ICD-based retrieval
gain_reading: A zero-shot LLM-assisted workflow improved automated identification of cardiovascular events from EMRs, achieving the highest AUCs for stroke, MI and composite MACE in two cohorts compared to ICD codes, primary diagnosis and problem list methods.
gain_evidence: LLM achieved the highest AUC for stroke (0.920; 95% CI 0.881 to 0.958), MI (0.938; 95% CI 0.905 to 0.971) and composite MACE (0.880; 95% CI 0.854 to 0.907) | In Cohort 2, the LLM achieved the highest AUC for all evaluated outcomes: stroke (0.915; 95% CI 0.862 to 0.968), MI (0.928; 95% CI 0.839 to 1.000), HF (0.844; 95% CI 0.803 to 0.884) and composite MACE (0.862; 95% CI 0.829 to 0.895)
problem_reading: The LLM-assisted workflow did not consistently outperform ICD-based retrieval, with no statistically significant AUC difference in Cohort 1 and lower AUC than ICD codes for heart failure in that cohort.
problem_evidence: ICD-based retrieval had a higher AUC for HF (0.882; 95% CI 0.845 to 0.918 vs 0.873; 95% CI 0.831 to 0.914) | differences in AUC between the LLM and ICD methods were not statistically significant across outcomes
quick_read: A multisite retrospective validation study in a US tertiary health system compared four automated EMR retrieval methods to manual chart adjudication for ischaemic stroke/TIA, MI, HF exacerbation/hospitalisation and composite MACE in 2258 patients treated with immune checkpoint inhibitors and 1426 patients who underwent TAVR. The zero-shot LLM workflow achieved the highest AUCs for most outcomes, while ICD-based retrieval remained competitive.

The findings matter because accurate automated capture of cardiovascular events is critical for retrospective outcomes research and pharmacovigilance, and the results suggest LLM extraction can complement but not uniformly replace structured code-based methods. Uncertainty remains about generalizability beyond a single health system, performance across other populations and outcomes, and prospective implementation costs and workflow integration.
limitation: Retrospective validation limited to three sites within a single US tertiary health system with two specific cohorts, and performance was context-dependent with ICD-based retrieval remaining competitive for some outcomes.
tag: Dual reading
key_points: Retrospective diagnostic accuracy study compared four automated EMR retrieval methods against clinician manual chart adjudication as reference standard. | Two adult cohorts from three sites within a single US tertiary health system: 2258 patients treated with immune checkpoint inhibitors and 1426 patients who underwent transcatheter aortic valve replacement. | Outcomes evaluated were ischaemic stroke or transient ischaemic attack, myocardial infarction, heart failure exacerbation or hospitalisation, and composite MACE. | AUC, sensitivity, specificity and net reclassification improvement were assessed for ICD codes, primary diagnosis, problem list and zero-shot LLM workflow.
rundown: Researchers compared ICD codes, primary diagnosis, problem list, and a zero-shot LLM workflow against clinician manual chart adjudication in 2258 immune checkpoint inhibitor patients and 1426 TAVR patients across three sites.

LLM had highest AUCs for most outcomes, including 0.920 for stroke and 0.938 for MI in Cohort 1 and 0.915 for stroke and 0.928 for MI in Cohort 2, but ICD codes had higher AUC for HF in Cohort 1 and differences were not statistically significant in that cohort, while LLM was significantly higher for stroke and composite MACE in Cohort 2.
sources:
- peer_reviewed | BMJ Open | https://doi.org/10.1136/bmjopen-2025-116133 | 2026-08-13
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
dabd8d5645fff82d2ba9069f38f967c29fd3ff9639ae0e05e90bb5c1bae64ba3
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0790 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.