TruaceTracing the truth around AIMonday, July 20, 2026
Health·The Trace·Model-prefilled trace·Published 2026-07-20

LLM-generated pulmonary tuberculosis Q&A responses evaluated for patient-education quality and readability

Source article: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability

Background: Pulmonary tuberculosis (TB) is a chronic infectious disease that burdens patients and public health systems. Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust. Large language models (LLMs) show promise for TB health education, but systematic evaluation is lacking. Objective: To evaluate five large language models in pulmonary tuberculosis Q&A (Question and Answer) scenarios and examine the effects of different large langua…

TRV-2026-0326Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 72The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 68The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability

"Classroom students png sticker illustration" is marked with CC0 1.0. To view the terms, visit https://creativecommons.org/publicdomain/zero/1.0/.

The quick read

From October 5 to 11, 2025, researchers tested five large language models on 20 pulmonary tuberculosis questions spanning five themes, generating 100 responses and rating them with C-PEMAT-P, GQS, and seven readability measures. GPT-5 ranked highest on C-PEMAT-P followed by Doubao, GQS was similar across models, and models differed significantly on several readability indices.

The findings matter because pulmonary TB burdens patients and public health systems and uneven information can reduce understanding and adherence, making AI-assisted education potentially useful but sensitive to model choice. It remains uncertain how these lab-based quality and readability scores translate to actual patient comprehension, trust, and outcomes, which the authors note requires more models, diseases, and patient-reported data.

Main points
  • Cross-sectional study conducted October 5 to October 11, 2025 using 20 pulmonary TB questions across five themes developed with a respiratory physician.
  • Five models tested: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT, producing 100 text responses evaluated with C-PEMAT-P, GQS, and seven readability indices including ARI and FRES.
  • GQS scores were similar across models while C-PEMAT-P favored GPT-5, and statistical analysis included One-way ANOVA, Kruskal-Wallis, and correlation analysis.
Gain

In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.

Problem

The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.

The rundown

Researchers entered 20 categorized TB questions into Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT to generate 100 responses, then scored quality with patient-education suitability and Global Quality Score and readability with seven indices including Automated Readability Index and Flesch Reading Ease Score.

Analysis found quality indicators were modestly associated with readability while readability indices were strongly intercorrelated, and themes had limited effects compared to model type as a key determinant of text quality.

What this doesn’t fix

Evaluation limited to five models and 20 physician-assisted questions without patient-reported outcomes, requiring broader testing across models, diseases, and real patient use.

Sources

Reader signal

How should this claim be treated?

The debate