LLM responses to common periodontal patient questions
Source article: Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions
Abstract: Background Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aimed to comparatively evaluate the accuracy and usability of widely used LLMs in responding to frequently asked periodontal questions. Methods In this analytical-comparative study, 15 commonly asked…
Negative state: both sides are scored from claims and sources, not community votes.

"09-8018-23" by NavyMedicine is marked with Public Domain Mark 1.0. To view the terms, visit https://creativecommons.org/publicdomain/mark/1.0/.
On 2026-09-16, a peer-reviewed comparative study reported testing 10 large language models on 15 common periodontal questions in English and Persian, with blinded ratings by two board-certified periodontists across six criteria. Performance varied significantly by model architecture, language, and question type, with the most advanced configuration providing highly accurate and well-referenced answers.
The findings matter because patients increasingly use LLMs for initial dental guidance, yet the study documents that outputs remain significantly better in English than Persian and often too technical for individuals with limited scientific literacy, lacking nuance for personalized risk assessment. This leaves uncertainty about safe integration into care without clinician filtering and language-specific optimization.
- 15 commonly asked periodontal questions from real patient encounters were administered to ten LLMs in both English and Persian.
- Responses were evaluated independently and blindly by two board-certified periodontists across six criteria including scientific accuracy and usability for limited literacy.
- Statistical analysis found significant differences among LLMs with chi-square 294.78, p reported, influenced by model architecture, language, and question type.
Advanced LLMs like ChatGPT-4.5 can provide accurate, comprehensive periodontal information to patients seeking quick answers to common gum disease questions.
LLMs answering periodontal questions perform significantly worse in Persian than English and often produce explanations too technical for average patients, lacking clinical nuance for personalized risk assessment.
The rundown
Researchers tested ten LLMs on 15 real patient periodontal questions in English and Persian, rating each response on a 5-point Likert scale for correlation, adequacy, comprehensiveness, clarity, usability for limited scientific literacy, and scientific accuracy.
Results showed variable performance by model and language, with deep-search ChatGPT-4.5 rated highest for accuracy and referencing, but overall tools were judged too technical for lay users and lacking nuance for complex treatment planning, leading authors to recommend use only as complementary to professional advice.
Study limited to 15 selected periodontal questions and 10 LLMs, with evaluation by two periodontists, and found language and literacy gaps requiring further optimization.
Sources
- Peer-reviewedClinical Advances in Periodontics2026-09-16
How should this claim be treated?
ace
The debate