TruaceTracing the truth around AISunday, September 20, 2026
TRV-2026-1131Version 1 · Certified

Written 2026-09-18 06:55:33 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1131
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-18T06:55:33.968148Z
status: published
lens: trace
sector: health
headline: Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions
dek: Background Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aimed to comparatively evaluate the accuracy and usability of widely used LLMs in responding to frequently asked periodontal questions. Methods In this analytical-comparative study, 15 commonly asked…
gain_title: Advanced LLMs like ChatGPT-4.5 can provide accurate, comprehensive periodontal information to patients seeking quick answers to common gum disease questions.
problem_title: LLMs answering periodontal questions perform significantly worse in Persian than English and often produce explanations too technical for average patients, lacking clinical nuance for personalized risk assessment.
trace_subject: LLM responses to common periodontal patient questions
gain_reading: Advanced LLMs like ChatGPT-4.5 can provide accurate, comprehensive periodontal information to patients seeking quick answers to common gum disease questions.
gain_evidence: provide accurate, high-quality information | highly accurate, comprehensive, and well-referenced responses
problem_reading: LLMs answering periodontal questions perform significantly worse in Persian than English and often produce explanations too technical for average patients, lacking clinical nuance for personalized risk assessment.
problem_evidence: significantly better in English than in Persian | too technical and hard for the average person to understand | often lack the clinical nuance required for personalized risk assessment
quick_read: On 2026-09-16, a peer-reviewed comparative study reported testing 10 large language models on 15 common periodontal questions in English and Persian, with blinded ratings by two board-certified periodontists across six criteria. Performance varied significantly by model architecture, language, and question type, with the most advanced configuration providing highly accurate and well-referenced answers.

The findings matter because patients increasingly use LLMs for initial dental guidance, yet the study documents that outputs remain significantly better in English than Persian and often too technical for individuals with limited scientific literacy, lacking nuance for personalized risk assessment. This leaves uncertainty about safe integration into care without clinician filtering and language-specific optimization.
limitation: Study limited to 15 selected periodontal questions and 10 LLMs, with evaluation by two periodontists, and found language and literacy gaps requiring further optimization.
tag: Dual reading
key_points: 15 commonly asked periodontal questions from real patient encounters were administered to ten LLMs in both English and Persian. | Responses were evaluated independently and blindly by two board-certified periodontists across six criteria including scientific accuracy and usability for limited literacy. | Statistical analysis found significant differences among LLMs with chi-square 294.78, p reported, influenced by model architecture, language, and question type.
rundown: Researchers tested ten LLMs on 15 real patient periodontal questions in English and Persian, rating each response on a 5-point Likert scale for correlation, adequacy, comprehensiveness, clarity, usability for limited scientific literacy, and scientific accuracy.

Results showed variable performance by model and language, with deep-search ChatGPT-4.5 rated highest for accuracy and referencing, but overall tools were judged too technical for lay users and lacking nuance for complex treatment planning, leading authors to recommend use only as complementary to professional advice.
sources:
- peer_reviewed | Clinical Advances in Periodontics | https://doi.org/10.1002/cap.70105 | 2026-09-16
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
6661969c0897f80078323850ae844824b0dbc4a941561bd0e2fdf3a6f924dfd8
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1131 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.