Evaluating the Accuracy, Empathy, and Readability of Generative AI Versus Registered Nurses in Discharge Planning: A Vignette-Based Study
Abstract: Aim To compare the multidimensional performance of discharge instructions generated by generative AI (GPT-4) versus those created by clinical registered nurses across three dimensions-accuracy, empathy and readability-and to explore the impact of patient. Design A prospective, double-blind, vignette-based cross-sectional study. Methods Five standardized multidisciplinary discharge scenarios were constructed. Discharge instructions were generated independently by five registered nurses and GPT-4. Fifteen clinical…
US Navy 060608-N-3714J-085 Navy Lt. Tara Collins, of Le Grange, Ky., gives a group of nursing students from the Jolo Norte Dame Nursing School a tour of the U.S. Military Sealift Command (MSC) Hospital ship USNS Mercy (T-AH 19) by U.S. Navy photo. Public domain
A prospective double-blind vignette study compared discharge instructions created by GPT-4 and by five registered nurses across five standardized scenarios. Fifteen experts rated accuracy and 38 patients rated empathy and readability, with NLP analysis of text complexity, using paired tests and a generalized linear mixed model.
The comparison matters because discharge education directly affects medication safety and patient understanding. While AI improved information completeness, the observed safety risks, lower empathy, higher linguistic complexity, and reduced acceptance among older and less-educated patients indicate it cannot yet replace nurses and requires mandatory nurse review and health-literacy-adapted deployment.
- Prospective double-blind vignette study compared GPT-4 vs 5 registered nurses across 5 multidisciplinary discharge scenarios.
- 15 clinical experts blindly rated accuracy; 38 patients blindly rated empathy and readability.
- Objective NLP analysis found AI texts had higher syntactic complexity and terminology density than nurse texts.
- Generalized linear mixed model linked advancing age and lower educational attainment to reduced acceptance of AI-generated texts.
In the same vignettes, experts identified safety risks in AI-generated discharge instructions not found in nurse texts, and nurses significantly outperformed AI on empathy and readability.
The rundown
Five standardized multidisciplinary discharge scenarios were used; instructions were generated independently by five registered nurses and GPT-4. Fifteen clinical experts conducted blinded accuracy assessments, while 38 patients conducted blinded empathy and readability assessments, with objective text features extracted via natural language processing.
Results showed AI led on comprehensiveness but introduced clinically unsafe content, while nurses led on empathy and readability. NLP confirmed higher syntactic complexity and terminology density in AI texts, and modeling showed older and less-educated patients were less likely to accept AI-generated texts, supporting a human-in-the-loop deployment model.
Sources
- Peer-reviewedNursing Open2026-09-01
How should this claim be treated?
ace
The debate