Consistency and accuracy of different artificial intelligence models in evaluating longitudinal cracks and fractures of teeth
Large language models such as ChatGPT and Google Gemini have attracted increasing attention in dentistry due to their potential to support clinical decision-making; however, their reliability in responding to endodontic questions remains unclear. This study compared the accuracy, consistency, and temporal stability of three artificial intelligence models (ChatGPT-3.5, ChatGPT-4o, and Google Gemini) when answering dichotomous clinical questions on vertical root fractures and tooth cracks using expert responses as…
ChatGPT-4o answered expert-validated true/false questions on vertical root fractures and tooth cracks with 86.1% accuracy, outperforming ChatGPT-3.5 and Gemini in a 5,400-response repeated test.
The same models showed very low inter-model agreement and significant temporal instability, with Gemini performing worse in evening sessions, indicating they cannot replace clinician judgment for vertical root fracture decisions.
Findings are limited to dichotomous true/false questions derived from a single position statement, with very low inter-model agreement and significant temporal variability, supporting use only as supportive tools.
Evidence
- Peer-reviewedOdontology2026-10-04
How should this claim be treated?
Truvace Impact Record TRV-2026-1305, v1: “Consistency and accuracy of different artificial intelligence models in evaluating longitudinal cracks and fractures of teeth.” Truvace, 2026-10-07. /record/TRV-2026-1305 (accessed at citation time). sha256 b2a6b71ce0f34b09…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1305 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace