Longitudinal symptom severity tracking in vagus nerve stimulation patients: a 2-stage LLM-based pipeline with explainable AI
Objectives To produce the first structured longitudinal dataset of symptom severity trajectories in vagus nerve stimulation (VNS) patients using a 2-stage large language model pipeline for automated symptom severity extraction from clinical notes, as part of the NIH-funded research evaluating vagal excitations and anatomical linkages (U54AT012307). Materials and methods We analyzed 1427 annotated clinical notes from 95 patients across 5 corpora. LLaMA models (3.2-1B, 3.2-3B, and 3.3-70B) were fine-tuned to class…

In brief
Researchers built a two-stage large language model pipeline to turn free-text clinical notes into structured symptom severity data for vagus nerve stimulation patients, part of NIH-funded work U54AT012307. Using 1427 annotated notes from 95 patients, LLaMA models classified symptoms and then graded severity, producing quarterly trajectories for 23 VNS patients compared to a 10-patient NSQIP surgical control.
The work matters because it moves VNS monitoring from population averages to individualized trajectories that could inform treatment-response prediction and device programming, but uncertainty remains about cross-domain reliability for low-prevalence psychiatric symptoms and about causal attribution given the small, unmatched control cohort and exploratory equity analyses.
Main points
- Analyzed 1427 annotated clinical notes from 95 patients across 5 corpora using LLaMA 3.2-1B, 3.2-3B, and 3.3-70B models.
- Stage 1 classified 13 symptoms as present, absent, or negated; stage 2 graded severity as indeterminate, mild, moderate, or severe.
- Validation corpus (VNS vs NSQIP) explained 29.4% of variance with no significant source-corpus effect, cited as support for cross-domain generalizability.
- Longitudinal tracking in 23 VNS patients versus 10-patient surgical control cohort showed median 0.5 quarters to meaningful seizure improvement.
The gain
Fine-tuned LLaMA models extracted symptom presence and severity from free-text notes with high F1 scores, enabling the first structured quarterly trajectories for VNS patients.
The rundown
The team fine-tuned three LLaMA models to first detect 13 symptoms and then assign severity levels, evaluating performance across five corpora including VNS and NSQIP notes. They report that validation corpus was the dominant variance factor at 29.4% with no significant source-corpus effect.
From quarterly trajectories in 23 VNS patients, the authors observed 50% responders, 25% non-responders, and 25% worsened, with median 0.5 quarters to improvement, while noting the absence of post-surgical decline in the unmatched 10-patient NSQIP cohort does not confirm VNS specificity.
Sources
- Peer-reviewedJAMIA Open2026-10-01
ace
The debate