TruaceTracing the truth around AISaturday, September 5, 2026
Health·G Space·Evidence-backed gain·Published 2026-09-05

Accuracy of General-Use Multimodal AI Platforms for Pell and Gregory Classification of Impacted Mandibular Third Molars

Abstract: Purpose The purpose of this study was to evaluate the performance of two general-use artificial intelligence models, ChatGPT and Grok, in classifying impacted mandibular third molars using the Pell and Gregory system on panoramic radiographs, compared with a resident consensus reference standard. Materials and methods One hundred panoramic radiographic images of impacted mandibular third molars were independently classified by two blinded resident reviewers using the Pell and Gregory classification system. Resid…

TRV-2026-0985Peer-reviewedPermanent record — cite & verify
Accuracy of General-Use Multimodal AI Platforms for Pell and Gregory Classification of Impacted Mandibular Third Molars

Role of Pharmaceutical Personnel In Tumor Board Closing the Gap of Cancer Care by Kauke Zimbwe. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

The quick read

Purpose The purpose of this study was to evaluate the performance of two general-use artificial intelligence models, ChatGPT and Grok, in classifying impacted mandibular third molars using the Pell and Gregory system on panoramic radiographs, compared with a resident consensus reference standard. Materials and methods One hundred panoramic radiographic images of impacted mandibular third molars were independently classified by two blinded resident reviewers using the Pell and Gregory classification system.

Resident consensus was defined as exact agreement between both reviewers; the 94 concordant classifications were confirmed by a board-certified oral and maxillofacial radiologist (the second author), blinded to the AI outputs, and served as the reference standard. Cases without consensus were excluded from the AI accuracy analysis.

Main points
  • Purpose The purpose of this study was to evaluate the performance of two general-use artificial intelligence models, ChatGPT and Grok, in classifying impacted mandibular third molars using the Pell and Gregory system on panoramic radiographs, compared with a resident consensus reference standard.
  • Materials and methods One hundred panoramic radiographic images of impacted mandibular third molars were independently classified by two blinded resident reviewers using the Pell and Gregory classification system.
  • Resident consensus was defined as exact agreement between both reviewers; the 94 concordant classifications were confirmed by a board-certified oral and maxillofacial radiologist (the second author), blinded to the AI outputs, and served as the reference standard.
Gain

Cases without consensus were excluded from the AI accuracy analysis. AI accuracy was calculated as the proportion of correct classifications among consensus cases, and the two models were compared using McNemar's test with continuity correction.

Sources

Reader signal

How should this claim be treated?

The debate