TRV-2026-0315Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0315 version: 1 kind: certified reason: Certified into the record timestamp: 2026-07-20T08:46:27.517029Z status: published lens: trace sector: health headline: Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability dek: Objective To evaluate red-flag recognition, clinical safety, and the quality of patient actionability in responses generated by artificial intelligence (AI) chatbots to patient questions about nipple discharge. Methods This guideline-informed cross-sectional evaluation was conducted to assess the performance of AI chatbots in simulated nipple discharge consultations. A total of 36 English-language simulated patient questions were developed on the basis of clinical guidelines and real-world consultation scenarios… gain_title: In simulated consultations about nipple discharge, AI chatbots recognized most clinical warning features and were rated safe in most responses. problem_title: A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations. trace_subject: triage safety of AI chatbots answering patient questions about nipple discharge gain_reading: In simulated consultations about nipple discharge, AI chatbots recognized most clinical warning features and were rated safe in most responses. gain_evidence: 189 (87.5 %) were rated as safe problem_reading: A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations. problem_evidence: 8 (3.7 %; 95 % confidence interval [CI], 1.4 %-6.9 %) met the primary outcome of potentially misleading/unsafe responses | Potentially misleading responses were most commonly attributable to missed red-flag features (4 of 8, 50.0 %) and poor actionability or insufficient action-oriented recommendations (3 of 8, 37.5 %) | some responses demonstrated incomplete red-flag safety-netting and lacked specific recommendations for subsequent action quick_read: Researchers evaluated six AI chatbots using 36 simulated English-language patient questions about nipple discharge, generating 216 first-turn responses scored against clinical guidelines. By the July 2026 publication date, 87.5% of responses were rated safe, 8.8% had minor omissions, and 3.7% were potentially misleading, with an overall red-flag recognition rate of 90.6%. The findings matter because patients may use chatbots for initial symptom information before seeking care, and missed warning features or vague next steps could delay evaluation. The authors conclude chatbots may serve as preliminary sources but cannot replace professional evaluation, and call for structured red-flag screening and explicit triage recommendations, while noting uncertainty due to simulated data and few unsafe events. limitation: Findings are limited to simulated English-language first-turn consultations with few unsafe events, so model comparisons are uncertain. tag: Model-prefilled trace key_points: Evaluation used 36 English-language simulated patient questions covering 6 modules of nipple discharge inquiries. | Six commonly used AI chatbots produced 216 first-turn responses evaluated by two independent reviewers with a third adjudicator. | Primary outcome was proportion of potentially misleading/unsafe responses; secondary measures included red-flag recognition rate, patient actionability score, guideline concordance, and DISCERN score. rundown: The study tested six chatbots on 36 simulated questions, producing 216 responses scored against a guideline-informed reference standard. Of those, 19 contained minor omissions and 8 met the primary outcome of potentially misleading/unsafe, with all 8 classified as potentially misleading and no potentially unsafe responses identified. Exploratory analysis found no clear statistical difference in the proportion of potentially misleading/unsafe responses across models when accounting for repeated-response structure, but observed between-model differences in red-flag recognition rate, patient actionability score, and total DISCERN score. sources: - peer_reviewed | International Journal of Medical Informatics | https://doi.org/10.1016/j.ijmedinf.2026.106603 | 2026-07-08 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- e5e02f60cdec7b9a4954302b13703ec99d04f3ddee706aa941927374f68e6899
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0315 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace