Multidisciplinary dental treatment planning by artificial intelligence: A comparative evaluation

Purpose The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs. Material and methods Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt…

Multidisciplinary dental treatment planning by artificial intelligence: A comparative evaluation
"22-0008-010 Camp Lejeune" by NavyMedicine is marked with Public Domain Mark 1.0. To view the terms, visit https://creativecommons.org/publicdomain/mark/1.0/.

In brief

AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case). Results Substantial inter-evaluator agreement was observed (κ = 0.74; 95% CI: 0.68-0.80), indicating consistent application of the scoring rubric.

Main points

  1. Purpose The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs.
  2. Material and methods Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt.
  3. A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist).

The problem

OpenEvidence and ChatGPT-5 demonstrated greater agreement with the expert reference standard, although clinically significant diagnostic errors were observed across all evaluated systems.

The rundown

A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist). AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case).

Sources

  1. Peer-reviewedJournal of Prosthodontics2026-08-21

The debate