Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP
Background National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction. Study design Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2…
Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p Conclusions In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables.
Evidence
- Peer-reviewedJournal of the American College of Surgeons2026-09-18
How should this claim be treated?
Truvace Impact Record TRV-2026-1141, v1: “Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP.” Truvace, 2026-09-19. /record/TRV-2026-1141 (accessed at citation time). sha256 05bacd8da620b511…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1141 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace