TruaceTracing the truth around AIFriday, September 18, 2026
Health·The Trace·Dual reading·Published 2026-09-18

LLM responses to common periodontal patient questions

Source article: Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions

Abstract: Background Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aimed to comparatively evaluate the accuracy and usability of widely used LLMs in responding to frequently asked periodontal questions. Methods In this analytical-comparative study, 15 commonly asked…

TRV-2026-1131Peer-reviewedPermanent record — cite & verify
Trace impact reading

Negative state: both sides are scored from claims and sources, not community votes.

P 73The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 67The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions

"09-8018-23" by NavyMedicine is marked with Public Domain Mark 1.0. To view the terms, visit https://creativecommons.org/publicdomain/mark/1.0/.

The quick read

On 2026-09-16, a peer-reviewed comparative study reported testing 10 large language models on 15 common periodontal questions in English and Persian, with blinded ratings by two board-certified periodontists across six criteria. Performance varied significantly by model architecture, language, and question type, with the most advanced configuration providing highly accurate and well-referenced answers.

The findings matter because patients increasingly use LLMs for initial dental guidance, yet the study documents that outputs remain significantly better in English than Persian and often too technical for individuals with limited scientific literacy, lacking nuance for personalized risk assessment. This leaves uncertainty about safe integration into care without clinician filtering and language-specific optimization.

Main points
  • 15 commonly asked periodontal questions from real patient encounters were administered to ten LLMs in both English and Persian.
  • Responses were evaluated independently and blindly by two board-certified periodontists across six criteria including scientific accuracy and usability for limited literacy.
  • Statistical analysis found significant differences among LLMs with chi-square 294.78, p reported, influenced by model architecture, language, and question type.
Gain

Advanced LLMs like ChatGPT-4.5 can provide accurate, comprehensive periodontal information to patients seeking quick answers to common gum disease questions.

Problem

LLMs answering periodontal questions perform significantly worse in Persian than English and often produce explanations too technical for average patients, lacking clinical nuance for personalized risk assessment.

The rundown

Researchers tested ten LLMs on 15 real patient periodontal questions in English and Persian, rating each response on a 5-point Likert scale for correlation, adequacy, comprehensiveness, clarity, usability for limited scientific literacy, and scientific accuracy.

Results showed variable performance by model and language, with deep-search ChatGPT-4.5 rated highest for accuracy and referencing, but overall tools were judged too technical for lay users and lacking nuance for complex treatment planning, leading authors to recommend use only as complementary to professional advice.

What this doesn’t fix

Study limited to 15 selected periodontal questions and 10 LLMs, with evaluation by two periodontists, and found language and literacy gaps requiring further optimization.

Sources

Reader signal

How should this claim be treated?

The debate