triage safety of AI chatbots answering patient questions about nipple discharge
Source article: Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability
Objective To evaluate red-flag recognition, clinical safety, and the quality of patient actionability in responses generated by artificial intelligence (AI) chatbots to patient questions about nipple discharge. Methods This guideline-informed cross-sectional evaluation was conducted to assess the performance of AI chatbots in simulated nipple discharge consultations. A total of 36 English-language simulated patient questions were developed on the basis of clinical guidelines and real-world consultation scenarios…
Negative state: both sides are scored from claims and sources, not community votes.

"Tomy Chatbot" by Latente 囧 www.latente.it is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.
Researchers evaluated six AI chatbots using 36 simulated English-language patient questions about nipple discharge, generating 216 first-turn responses scored against clinical guidelines. By the July 2026 publication date, 87.5% of responses were rated safe, 8.8% had minor omissions, and 3.7% were potentially misleading, with an overall red-flag recognition rate of 90.6%.
The findings matter because patients may use chatbots for initial symptom information before seeking care, and missed warning features or vague next steps could delay evaluation. The authors conclude chatbots may serve as preliminary sources but cannot replace professional evaluation, and call for structured red-flag screening and explicit triage recommendations, while noting uncertainty due to simulated data and few unsafe events.
- Evaluation used 36 English-language simulated patient questions covering 6 modules of nipple discharge inquiries.
- Six commonly used AI chatbots produced 216 first-turn responses evaluated by two independent reviewers with a third adjudicator.
- Primary outcome was proportion of potentially misleading/unsafe responses; secondary measures included red-flag recognition rate, patient actionability score, guideline concordance, and DISCERN score.
In simulated consultations about nipple discharge, AI chatbots recognized most clinical warning features and were rated safe in most responses.
A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations.
The rundown
The study tested six chatbots on 36 simulated questions, producing 216 responses scored against a guideline-informed reference standard. Of those, 19 contained minor omissions and 8 met the primary outcome of potentially misleading/unsafe, with all 8 classified as potentially misleading and no potentially unsafe responses identified.
Exploratory analysis found no clear statistical difference in the proportion of potentially misleading/unsafe responses across models when accounting for repeated-response structure, but observed between-model differences in red-flag recognition rate, patient actionability score, and total DISCERN score.
Findings are limited to simulated English-language first-turn consultations with few unsafe events, so model comparisons are uncertain.
Sources
- Peer-reviewedInternational Journal of Medical Informatics2026-07-08
How should this claim be treated?
ace
The debate