TruaceTracing the truth around AISaturday, September 19, 2026
Health·P Space·Evidence-backed problem·Published 2026-09-19

Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP

Abstract: Background National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction. Study design Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2…

TRV-2026-1141Peer-reviewedPermanent record — cite & verify
Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP

Abstraction in the INTEL iAPX-432 prototype systems implementation language by MacLennan, Bruce J.. Public domain

The quick read

Background National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction.

Study design Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. Overall accuracy of LLM and human abstraction was compared using McNemar's and Chi-square tests.

Main points
  • Background National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling.
  • This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction.
  • Study design Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables.
Problem

Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p Conclusions In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables.

The rundown

Study design Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. A faculty plastic surgeon established the reference standard.

Sources

Reader signal

How should this claim be treated?

The debate