TRV-2026-0326Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0326 version: 1 kind: certified reason: Certified into the record timestamp: 2026-07-20T08:46:27.995478Z status: published lens: trace sector: health headline: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability dek: Background: Pulmonary tuberculosis (TB) is a chronic infectious disease that burdens patients and public health systems. Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust. Large language models (LLMs) show promise for TB health education, but systematic evaluation is lacking. Objective: To evaluate five large language models in pulmonary tuberculosis Q&A (Question and Answer) scenarios and examine the effects of different large langua… gain_title: In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs. problem_title: The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence. trace_subject: LLM-generated pulmonary tuberculosis Q&A responses evaluated for patient-education quality and readability gain_reading: In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs. gain_evidence: GPT-5 achieved the highest C-PEMAT-P scores, followed by Doubao problem_reading: The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence. problem_evidence: Models showed significant differences on several readability indices | Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust quick_read: From October 5 to 11, 2025, researchers tested five large language models on 20 pulmonary tuberculosis questions spanning five themes, generating 100 responses and rating them with C-PEMAT-P, GQS, and seven readability measures. GPT-5 ranked highest on C-PEMAT-P followed by Doubao, GQS was similar across models, and models differed significantly on several readability indices. The findings matter because pulmonary TB burdens patients and public health systems and uneven information can reduce understanding and adherence, making AI-assisted education potentially useful but sensitive to model choice. It remains uncertain how these lab-based quality and readability scores translate to actual patient comprehension, trust, and outcomes, which the authors note requires more models, diseases, and patient-reported data. limitation: Evaluation limited to five models and 20 physician-assisted questions without patient-reported outcomes, requiring broader testing across models, diseases, and real patient use. tag: Model-prefilled trace key_points: Cross-sectional study conducted October 5 to October 11, 2025 using 20 pulmonary TB questions across five themes developed with a respiratory physician. | Five models tested: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT, producing 100 text responses evaluated with C-PEMAT-P, GQS, and seven readability indices including ARI and FRES. | GQS scores were similar across models while C-PEMAT-P favored GPT-5, and statistical analysis included One-way ANOVA, Kruskal-Wallis, and correlation analysis. rundown: Researchers entered 20 categorized TB questions into Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT to generate 100 responses, then scored quality with patient-education suitability and Global Quality Score and readability with seven indices including Automated Readability Index and Flesch Reading Ease Score. Analysis found quality indicators were modestly associated with readability while readability indices were strongly intercorrelated, and themes had limited effects compared to model type as a key determinant of text quality. sources: - peer_reviewed | DIGITAL HEALTH | https://doi.org/10.1177/20552076261467813 | 2026-01-01 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- 397115fc536c270e774f86078ea8ac57dfbc02e95ad57281b2c1badc71380d4e
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0326 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace