TruaceTracing the truth around AITuesday, July 21, 2026
TRV-2026-0326Version 1 · Certified

Written 2026-07-20 08:46:27 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0326
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-20T08:46:27.995478Z
status: published
lens: trace
sector: health
headline: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability
dek: Background: Pulmonary tuberculosis (TB) is a chronic infectious disease that burdens patients and public health systems. Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust. Large language models (LLMs) show promise for TB health education, but systematic evaluation is lacking. Objective: To evaluate five large language models in pulmonary tuberculosis Q&A (Question and Answer) scenarios and examine the effects of different large langua…
gain_title: In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.
problem_title: The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.
trace_subject: LLM-generated pulmonary tuberculosis Q&A responses evaluated for patient-education quality and readability
gain_reading: In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.
gain_evidence: GPT-5 achieved the highest C-PEMAT-P scores, followed by Doubao
problem_reading: The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.
problem_evidence: Models showed significant differences on several readability indices | Limited reach of traditional education and uneven online information may undermine patients' understanding, adherence, and trust
quick_read: From October 5 to 11, 2025, researchers tested five large language models on 20 pulmonary tuberculosis questions spanning five themes, generating 100 responses and rating them with C-PEMAT-P, GQS, and seven readability measures. GPT-5 ranked highest on C-PEMAT-P followed by Doubao, GQS was similar across models, and models differed significantly on several readability indices.

The findings matter because pulmonary TB burdens patients and public health systems and uneven information can reduce understanding and adherence, making AI-assisted education potentially useful but sensitive to model choice. It remains uncertain how these lab-based quality and readability scores translate to actual patient comprehension, trust, and outcomes, which the authors note requires more models, diseases, and patient-reported data.
limitation: Evaluation limited to five models and 20 physician-assisted questions without patient-reported outcomes, requiring broader testing across models, diseases, and real patient use.
tag: Model-prefilled trace
key_points: Cross-sectional study conducted October 5 to October 11, 2025 using 20 pulmonary TB questions across five themes developed with a respiratory physician. | Five models tested: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT, producing 100 text responses evaluated with C-PEMAT-P, GQS, and seven readability indices including ARI and FRES. | GQS scores were similar across models while C-PEMAT-P favored GPT-5, and statistical analysis included One-way ANOVA, Kruskal-Wallis, and correlation analysis.
rundown: Researchers entered 20 categorized TB questions into Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and ChatGPT to generate 100 responses, then scored quality with patient-education suitability and Global Quality Score and readability with seven indices including Automated Readability Index and Flesch Reading Ease Score.

Analysis found quality indicators were modestly associated with readability while readability indices were strongly intercorrelated, and themes had limited effects compared to model type as a key determinant of text quality.
sources:
- peer_reviewed | DIGITAL HEALTH | https://doi.org/10.1177/20552076261467813 | 2026-01-01
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
397115fc536c270e774f86078ea8ac57dfbc02e95ad57281b2c1badc71380d4e
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0326 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.