Multidisciplinary dental treatment planning by artificial intelligence: A comparative evaluation
Abstract: Purpose The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs. Material and methods Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt…

"22-0008-010 Camp Lejeune" by NavyMedicine is marked with Public Domain Mark 1.0. To view the terms, visit https://creativecommons.org/publicdomain/mark/1.0/.
AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case). Results Substantial inter-evaluator agreement was observed (κ = 0.74; 95% CI: 0.68-0.80), indicating consistent application of the scoring rubric.
- Purpose The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs.
- Material and methods Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt.
- A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist).
OpenEvidence and ChatGPT-5 demonstrated greater agreement with the expert reference standard, although clinically significant diagnostic errors were observed across all evaluated systems.
The rundown
A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist). AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case).
Sources
- Peer-reviewedJournal of Prosthodontics2026-08-21
How should this claim be treated?
ace
The debate