LLM-generated pulmonary tuberculosis Q&A responses evaluated for patient-education quality and readability
Source article: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability
Background: Pulmonary tuberculosis (TB) is a chronic infectious disease that burdens patients and public health systems. Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust. Large language models (LLMs) show promise for TB health education, but systematic evaluation is lacking. Objective: To evaluate five large language models in pulmonary tuberculosis Q&A (Question and Answer) scenarios and examine the effects of different large langua…
Contested: both sides are scored from claims and sources, not community votes.
"Classroom students png sticker illustration" is marked with CC0 1.0. To view the terms, visit https://creativecommons.org/publicdomain/zero/1.0/.
From October 5 to 11, 2025, researchers tested five large language models on 20 pulmonary tuberculosis questions spanning five themes, generating 100 responses and rating them with C-PEMAT-P, GQS, and seven readability measures. GPT-5 ranked highest on C-PEMAT-P followed by Doubao, GQS was similar across models, and models differed significantly on several readability indices.
The findings matter because pulmonary TB burdens patients and public health systems and uneven information can reduce understanding and adherence, making AI-assisted education potentially useful but sensitive to model choice. It remains uncertain how these lab-based quality and readability scores translate to actual patient comprehension, trust, and outcomes, which the authors note requires more models, diseases, and patient-reported data.
- Cross-sectional study conducted October 5 to October 11, 2025 using 20 pulmonary TB questions across five themes developed with a respiratory physician.
- Five models tested: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT, producing 100 text responses evaluated with C-PEMAT-P, GQS, and seven readability indices including ARI and FRES.
- GQS scores were similar across models while C-PEMAT-P favored GPT-5, and statistical analysis included One-way ANOVA, Kruskal-Wallis, and correlation analysis.
In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.
The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.
The rundown
Researchers entered 20 categorized TB questions into Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT to generate 100 responses, then scored quality with patient-education suitability and Global Quality Score and readability with seven indices including Automated Readability Index and Flesch Reading Ease Score.
Analysis found quality indicators were modestly associated with readability while readability indices were strongly intercorrelated, and themes had limited effects compared to model type as a key determinant of text quality.
Evaluation limited to five models and 20 physician-assisted questions without patient-reported outcomes, requiring broader testing across models, diseases, and real patient use.
Sources
- Peer-reviewedDIGITAL HEALTH2026-01-01
How should this claim be treated?
ace
The debate