quality and safety of guideline-grounded rheumatology chatbot responses for patient self-management

Source article: Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response chara…

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology
"MIHANOVICEVA 3 - rheumatology and physiotherapy" by Miroslav Vajdić is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.
Trace impact readingContested
P 70The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.

Both sides are scored from claims and sources, not community votes.

G 67The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.

In brief

Researchers co-developed ten disease-specific chatbots grounded in German rheumatology guidelines and deployed them through 13 rheumatology centres and six patient organisations from September 2025 to January 2026. They analyzed 6291 user questions, 2671 user ratings of responses, and 602 questionnaire responses, plus a six-dimensional LLM-as-a-judge assessment compared with rheumatologist ratings in subsets.

The deployment matters because it provides real-world evidence that patients will use and generally like guideline-grounded chatbots for persistent information needs, yet it also shows gaps in completeness and guideline fidelity. Uncertainty remains about whether positive ratings translate into improved self-management or clinical outcomes, and about how to independently validate safety given weak agreement between automated and physician judgments.

Main points

  1. Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations.
  2. Between September 2025 and January 2026, 6291 questions were recorded, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%).
  3. Of 2671 responses rated by users, 92.9% received a positive rating, while 84.6% of questionnaire respondents said the chatbot was easy to use.
  4. LLM-as-a-judge evaluation found 95.3% of answers completely safe and 79.1% completely correct, but only 45.0% fully adherent to guidelines with weak agreement with rheumatologist ratings.

The gain

Patients using guideline-grounded rheumatology chatbots reported high usability and satisfaction, with most answers rated safe and correct in real-world deployment.

The problem

Chatbots failed to answer 4.2% of interactions, received negative ratings for insufficient detail, and showed only 45% full guideline adherence with weak agreement between LLM and physician assessments.

The rundown

Between September 2025 and January 2026, ten German guideline-based chatbots were deployed via 13 centres and six patient organisations, logging 6291 questions. Questions were coded into 13 categories, with disease-specific, medication and monitoring, and diagnostics most frequent, and 263 interactions resulted in no answer.

User-rated responses were 92.9% positive among 2671 ratings, with insufficient detail cited in 65.8% of negative ratings. Among 602 questionnaire respondents, over 80% reported ease of use, understandability, and usefulness as an addition to education, while LLM-based evaluation rated 95.3% completely safe and 79.1% completely correct but only 45% fully guideline adherent.

What this doesn’t fix

Findings are limited by reliance on LLM-based quality ratings with weak agreement with physician assessment and lack of measured educational effectiveness or independent clinical safety validation.

Sources

  1. Peer-reviewedJournal of Medical Systems2026-10-05

The debate