AI models answering dichotomous clinical questions on vertical root fractures and tooth cracks
Source article: Consistency and accuracy of different artificial intelligence models in evaluating longitudinal cracks and fractures of teeth
Large language models such as ChatGPT and Google Gemini have attracted increasing attention in dentistry due to their potential to support clinical decision-making; however, their reliability in responding to endodontic questions remains unclear. This study compared the accuracy, consistency, and temporal stability of three artificial intelligence models (ChatGPT-3.5, ChatGPT-4o, and Google Gemini) when answering dichotomous clinical questions on vertical root fractures and tooth cracks using expert responses as…

Both sides are scored from claims and sources, not community votes.
G 65The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.In brief
A peer-reviewed study in Odontology compared ChatGPT-3.5, ChatGPT-4o, and Google Gemini on 60 expert-validated true/false questions about longitudinal tooth fractures, querying each model three times daily for 10 days for 5,400 total responses against expert reference answers.
The results matter because even the best model reached only 86.1% accuracy with very low agreement across models and measurable instability over time, leaving uncertainty about reliability in real clinical workflows and reinforcing that such tools should remain supportive rather than determinative in endodontics.
Main points
- Study used 60 validated true/false items based on European Society of Endodontology position statement on longitudinal tooth fractures, categorized into easy, moderate, and difficult levels.
- Each of three models was queried over 10 consecutive days at three daily time points, generating 5,400 responses total.
- Overall accuracy differed significantly among models (p < 0.001) and differences between models were evident only for moderate-difficulty questions (p < 0.001).
- Inter-model agreement was very low with Cohen's kappa (k) = 0.019 and temporal stability differed significantly among models (p < 0.001).
The gain
ChatGPT-4o answered expert-validated true/false questions on vertical root fractures and tooth cracks with 86.1% accuracy, outperforming ChatGPT-3.5 and Gemini in a 5,400-response repeated test.
The problem
The same models showed very low inter-model agreement and significant temporal instability, with Gemini performing worse in evening sessions, indicating they cannot replace clinician judgment for vertical root fracture decisions.
The rundown
Researchers tested ChatGPT-3.5, ChatGPT-4o, and Google Gemini on 60 validated true/false questions drawn from the European Society of Endodontology position statement, with item-level content validity index of 1.00, administered over 10 days at morning, afternoon, and evening time points.
Results showed ChatGPT-4o at 86.1% accuracy, ChatGPT-3.5 at 85.6%, and Gemini at 81.8%, with no morning-afternoon difference but significantly lower evening performance for Gemini, and very low inter-model agreement despite acceptable accuracy levels.
What this doesn’t fix
Findings are limited to dichotomous true/false questions derived from a single position statement, with very low inter-model agreement and significant temporal variability, supporting use only as supportive tools.
Sources
- Peer-reviewedOdontology2026-10-04
ace
The debate