LLM-VSP self-practice effect on undergraduate medical students' medical history-taking performance
Source article: Effect of Large Language Model-Powered Virtual Standardized Patients on History-Taking Among Undergraduate Medical Students: Propensity-Matched Cohort Study
Abstract: Background Medical history taking (MHT) is a foundational clinical competency for medical students; however, traditional training models using standardized patients face challenges such as resource constraints. Large language model-powered virtual standardized patients (LLM-VSPs) offer a safe, repeatable platform for self-directed practice with AI-automated feedback. Nevertheless, their effectiveness in authentic teaching environments and underlying learning mechanisms require further investigation. Objective Th…
Contested: both sides are scored from claims and sources, not community votes.
Hospital Universitari Doctor Peset, València 08 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
In a 2026 prospective cohort study, 168 third-year medical students were grouped by voluntary use of an LLM-powered virtual standardized patient system for extracurricular history-taking practice versus routine instruction alone. After propensity score matching to 40 pairs, the LLM-VSP group scored higher on an end-of-term Objective Structured Clinical Examination with real standardized patients.
The finding matters because virtual patients offer a repeatable, low-resource alternative to traditional standardized patient training, but the benefit was uneven. High-baseline students showed larger gains while medium and low-baseline students improved little, and practice counts alone did not predict outcomes, leaving uncertainty about how to support lower-proficiency learners and whether results generalize beyond this single-site, non-randomized design.
- Prospective cohort of 168 third-year medical students, 120 intervention using LLM-VSP and 48 control receiving routine instruction.
- Propensity score matching yielded 40 matched pairs to balance confounding factors before comparing outcomes.
- Primary outcome was end-of-term MHT performance assessed at an Objective Structured Clinical Examination station with real standardized patients.
- Exploratory analysis found practice behavior metrics were not independent predictors of final scores when constrained by baseline proficiency.
Undergraduate medical students who used LLM-powered virtual standardized patients as extracurricular self-practice achieved higher end-of-term history-taking performance at an OSCE with real standardized patients compared to routine instruction.
Students with medium and low baseline history-taking proficiency showed relatively limited score improvements from LLM-VSP self-practice, with practice frequency alone not independently predicting final performance.
The rundown
The study enrolled 168 third-year students after didactic instruction but before clinical practicum, with baseline MHT assessed via virtual patient examination and final performance measured at an OSCE station using real standardized patients.
After matching, baseline characteristics were balanced with standardized mean difference 2 =0.708, and robustness was checked with multiple linear regression, sensitivity analyses, and Rosenbaum bounds analyses.
Subgroup results showed high baseline students gained mean difference 7.04, 95% CI 3.84-10.25 in the overall sample, while medium and low baseline groups gained only 1.29-1.64, indicating differential benefit.
Effect may depend on baseline proficiency, with preliminary evidence of a cognitive threshold where only students with solid theoretical foundations benefit substantially, and assignment was based on voluntary participation requiring propensity matching rather than randomization.
Sources
- Peer-reviewedJMIR Medical Education2026-09-09
How should this claim be treated?
ace
The debate