TruaceTracing the truth around AISaturday, September 12, 2026
TRV-2026-1044Version 1 · Certified

Written 2026-09-10 06:05:12 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1044
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-10T06:05:12.367107Z
status: published
lens: trace
sector: health
headline: Effect of Large Language Model-Powered Virtual Standardized Patients on History-Taking Among Undergraduate Medical Students: Propensity-Matched Cohort Study
dek: Background Medical history taking (MHT) is a foundational clinical competency for medical students; however, traditional training models using standardized patients face challenges such as resource constraints. Large language model-powered virtual standardized patients (LLM-VSPs) offer a safe, repeatable platform for self-directed practice with AI-automated feedback. Nevertheless, their effectiveness in authentic teaching environments and underlying learning mechanisms require further investigation. Objective Th…
gain_title: Undergraduate medical students who used LLM-powered virtual standardized patients as extracurricular self-practice achieved higher end-of-term history-taking performance at an OSCE with real standardized patients compared to routine instruction.
problem_title: Students with medium and low baseline history-taking proficiency showed relatively limited score improvements from LLM-VSP self-practice, with practice frequency alone not independently predicting final performance.
trace_subject: LLM-VSP self-practice effect on undergraduate medical students' medical history-taking performance
gain_reading: Undergraduate medical students who used LLM-powered virtual standardized patients as extracurricular self-practice achieved higher end-of-term history-taking performance at an OSCE with real standardized patients compared to routine instruction.
gain_evidence: Introducing LLM-VSPs as a self-practice tool in diagnostics education may help improve undergraduate medical students' MHT performance.
problem_reading: Students with medium and low baseline history-taking proficiency showed relatively limited score improvements from LLM-VSP self-practice, with practice frequency alone not independently predicting final performance.
problem_evidence: improvements in medium and low baseline students were relatively limited (matched sample: mean difference 2.94-2.99; overall sample: mean difference 1.29-1.64) | mere practice behavior metrics were not independent predictors of final scores, potentially being constrained by baseline proficiency
quick_read: In a 2026 prospective cohort study, 168 third-year medical students were grouped by voluntary use of an LLM-powered virtual standardized patient system for extracurricular history-taking practice versus routine instruction alone. After propensity score matching to 40 pairs, the LLM-VSP group scored higher on an end-of-term Objective Structured Clinical Examination with real standardized patients.

The finding matters because virtual patients offer a repeatable, low-resource alternative to traditional standardized patient training, but the benefit was uneven. High-baseline students showed larger gains while medium and low-baseline students improved little, and practice counts alone did not predict outcomes, leaving uncertainty about how to support lower-proficiency learners and whether results generalize beyond this single-site, non-randomized design.
limitation: Effect may depend on baseline proficiency, with preliminary evidence of a cognitive threshold where only students with solid theoretical foundations benefit substantially, and assignment was based on voluntary participation requiring propensity matching rather than randomization.
tag: Dual reading
key_points: Prospective cohort of 168 third-year medical students, 120 intervention using LLM-VSP and 48 control receiving routine instruction. | Propensity score matching yielded 40 matched pairs to balance confounding factors before comparing outcomes. | Primary outcome was end-of-term MHT performance assessed at an Objective Structured Clinical Examination station with real standardized patients. | Exploratory analysis found practice behavior metrics were not independent predictors of final scores when constrained by baseline proficiency.
rundown: The study enrolled 168 third-year students after didactic instruction but before clinical practicum, with baseline MHT assessed via virtual patient examination and final performance measured at an OSCE station using real standardized patients.

After matching, baseline characteristics were balanced with standardized mean difference 2 =0.708, and robustness was checked with multiple linear regression, sensitivity analyses, and Rosenbaum bounds analyses.

Subgroup results showed high baseline students gained mean difference 7.04, 95% CI 3.84-10.25 in the overall sample, while medium and low baseline groups gained only 1.29-1.64, indicating differential benefit.
sources:
- peer_reviewed | JMIR Medical Education | https://doi.org/10.2196/92486 | 2026-09-09
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
75073b2f5a6fb37ccc61cca52907ee2089797e66b5b761b6e34779a6c21750b5
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1044 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.