TruaceTracing the truth around AIMonday, July 20, 2026
Health·The Trace·Model-prefilled trace·Published 2026-07-20

triage safety of AI chatbots answering patient questions about nipple discharge

Source article: Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability

Objective To evaluate red-flag recognition, clinical safety, and the quality of patient actionability in responses generated by artificial intelligence (AI) chatbots to patient questions about nipple discharge. Methods This guideline-informed cross-sectional evaluation was conducted to assess the performance of AI chatbots in simulated nipple discharge consultations. A total of 36 English-language simulated patient questions were developed on the basis of clinical guidelines and real-world consultation scenarios…

TRV-2026-0315Peer-reviewedPermanent record — cite & verify
Trace impact reading

Negative state: both sides are scored from claims and sources, not community votes.

P 70The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 65The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability

"Tomy Chatbot" by Latente 囧 www.latente.it is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.

The quick read

Researchers evaluated six AI chatbots using 36 simulated English-language patient questions about nipple discharge, generating 216 first-turn responses scored against clinical guidelines. By the July 2026 publication date, 87.5% of responses were rated safe, 8.8% had minor omissions, and 3.7% were potentially misleading, with an overall red-flag recognition rate of 90.6%.

The findings matter because patients may use chatbots for initial symptom information before seeking care, and missed warning features or vague next steps could delay evaluation. The authors conclude chatbots may serve as preliminary sources but cannot replace professional evaluation, and call for structured red-flag screening and explicit triage recommendations, while noting uncertainty due to simulated data and few unsafe events.

Main points
  • Evaluation used 36 English-language simulated patient questions covering 6 modules of nipple discharge inquiries.
  • Six commonly used AI chatbots produced 216 first-turn responses evaluated by two independent reviewers with a third adjudicator.
  • Primary outcome was proportion of potentially misleading/unsafe responses; secondary measures included red-flag recognition rate, patient actionability score, guideline concordance, and DISCERN score.
Gain

In simulated consultations about nipple discharge, AI chatbots recognized most clinical warning features and were rated safe in most responses.

Problem

A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations.

The rundown

The study tested six chatbots on 36 simulated questions, producing 216 responses scored against a guideline-informed reference standard. Of those, 19 contained minor omissions and 8 met the primary outcome of potentially misleading/unsafe, with all 8 classified as potentially misleading and no potentially unsafe responses identified.

Exploratory analysis found no clear statistical difference in the proportion of potentially misleading/unsafe responses across models when accounting for repeated-response structure, but observed between-model differences in red-flag recognition rate, patient actionability score, and total DISCERN score.

What this doesn’t fix

Findings are limited to simulated English-language first-turn consultations with few unsafe events, so model comparisons are uncertain.

Sources

Reader signal

How should this claim be treated?

The debate