TRV-2026-1292Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-1292 version: 1 kind: certified reason: Certified into the record timestamp: 2026-10-06T06:55:14.204132Z status: published lens: trace sector: health headline: Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology dek: Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response chara… gain_title: Patients using guideline-grounded rheumatology chatbots reported high usability and satisfaction, with most answers rated safe and correct in real-world deployment. problem_title: Chatbots failed to answer 4.2% of interactions, received negative ratings for insufficient detail, and showed only 45% full guideline adherence with weak agreement between LLM and physician assessments. trace_subject: quality and safety of guideline-grounded rheumatology chatbot responses for patient self-management gain_reading: Patients using guideline-grounded rheumatology chatbots reported high usability and satisfaction, with most answers rated safe and correct in real-world deployment. gain_evidence: 95.3% of answers were rated as completely safe and 79.1% as completely correct problem_reading: Chatbots failed to answer 4.2% of interactions, received negative ratings for insufficient detail, and showed only 45% full guideline adherence with weak agreement between LLM and physician assessments. problem_evidence: The chatbots were unable to answer in 263 interactions (4.2%) | Insufficient detail was the most common reason for negative ratings (125/190, 65.8%) | 45.0% rated as fully adherent; agreement with physician assessment was weak quick_read: Researchers co-developed ten disease-specific chatbots grounded in German rheumatology guidelines and deployed them through 13 rheumatology centres and six patient organisations from September 2025 to January 2026. They analyzed 6291 user questions, 2671 user ratings of responses, and 602 questionnaire responses, plus a six-dimensional LLM-as-a-judge assessment compared with rheumatologist ratings in subsets. The deployment matters because it provides real-world evidence that patients will use and generally like guideline-grounded chatbots for persistent information needs, yet it also shows gaps in completeness and guideline fidelity. Uncertainty remains about whether positive ratings translate into improved self-management or clinical outcomes, and about how to independently validate safety given weak agreement between automated and physician judgments. limitation: Findings are limited by reliance on LLM-based quality ratings with weak agreement with physician assessment and lack of measured educational effectiveness or independent clinical safety validation. tag: Dual reading key_points: Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. | Between September 2025 and January 2026, 6291 questions were recorded, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). | Of 2671 responses rated by users, 92.9% received a positive rating, while 84.6% of questionnaire respondents said the chatbot was easy to use. | LLM-as-a-judge evaluation found 95.3% of answers completely safe and 79.1% completely correct, but only 45.0% fully adherent to guidelines with weak agreement with rheumatologist ratings. rundown: Between September 2025 and January 2026, ten German guideline-based chatbots were deployed via 13 centres and six patient organisations, logging 6291 questions. Questions were coded into 13 categories, with disease-specific, medication and monitoring, and diagnostics most frequent, and 263 interactions resulted in no answer. User-rated responses were 92.9% positive among 2671 ratings, with insufficient detail cited in 65.8% of negative ratings. Among 602 questionnaire respondents, over 80% reported ease of use, understandability, and usefulness as an addition to education, while LLM-based evaluation rated 95.3% completely safe and 79.1% completely correct but only 45% fully guideline adherent. sources: - peer_reviewed | Journal of Medical Systems | https://doi.org/10.1007/s10916-026-02470-6 | 2026-10-05 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- ff14a54f27f6b02db8c22ff8446f024cf68d0518531a1815915268a8f3f7fa93
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1292 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace