Can Artificial Intelligence Match Human Expertise in Long-Term Periodontal Prognosis? A Comparative Accuracy Study
To compare the prognostic performance of an artificial intelligence (AI) model with that of experienced clinicians in predicting tooth loss over a 10-year period. An AI model trained on structured clinical and radiographic data was compared with 12 periodontists and 11 general dentists (GDs), who independently assigned prognostic scores (0-10 scale) to 300 teeth with known 10-year outcomes. AI and clinician performance were evaluated under a fixed threshold and group-specific optimal thresholds. Accuracy, sensit…

In brief
To compare the prognostic performance of an artificial intelligence (AI) model with that of experienced clinicians in predicting tooth loss over a 10-year period. An AI model trained on structured clinical and radiographic data was compared with 12 periodontists and 11 general dentists (GDs), who independently assigned prognostic scores (0-10 scale) to 300 teeth with known 10-year outcomes.
AI and clinician performance were evaluated under a fixed threshold and group-specific optimal thresholds. Using the fixed threshold (score > 5 = survival), clinicians achieved higher overall accuracy than AI (75.6% periodontists, 74.9% GDs, 69.2% AI; p < 0.05), with sensitivity low and comparable across groups (14.7%-22.7%).
Main points
- To compare the prognostic performance of an artificial intelligence (AI) model with that of experienced clinicians in predicting tooth loss over a 10-year period.
- An AI model trained on structured clinical and radiographic data was compared with 12 periodontists and 11 general dentists (GDs), who independently assigned prognostic scores (0-10 scale) to 300 teeth with known 10-year outcomes.
- AI and clinician performance were evaluated under a fixed threshold and group-specific optimal thresholds.
The gain
Using the fixed threshold (score > 5 = survival), clinicians achieved higher overall accuracy than AI (75.6% periodontists, 74.9% GDs, 69.2% AI; p < 0.05), with sensitivity low and comparable across groups (14.7%-22.7%).
The rundown
AI and clinician performance were evaluated under a fixed threshold and group-specific optimal thresholds. Accuracy, sensitivity, specificity, predictive values, area under the receiver operating characteristic curve and calibration were used for comparison.
Sources
- Peer-reviewedJournal of Clinical Periodontology2026-08-21
ace



The debate