TRV-2026-1305Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-1305 version: 1 kind: certified reason: Certified into the record timestamp: 2026-10-07T06:56:37.342542Z status: published lens: trace sector: health headline: Consistency and accuracy of different artificial intelligence models in evaluating longitudinal cracks and fractures of teeth dek: Large language models such as ChatGPT and Google Gemini have attracted increasing attention in dentistry due to their potential to support clinical decision-making; however, their reliability in responding to endodontic questions remains unclear. This study compared the accuracy, consistency, and temporal stability of three artificial intelligence models (ChatGPT-3.5, ChatGPT-4o, and Google Gemini) when answering dichotomous clinical questions on vertical root fractures and tooth cracks using expert responses as… gain_title: ChatGPT-4o answered expert-validated true/false questions on vertical root fractures and tooth cracks with 86.1% accuracy, outperforming ChatGPT-3.5 and Gemini in a 5,400-response repeated test. problem_title: The same models showed very low inter-model agreement and significant temporal instability, with Gemini performing worse in evening sessions, indicating they cannot replace clinician judgment for vertical root fracture decisions. trace_subject: AI models answering dichotomous clinical questions on vertical root fractures and tooth cracks gain_reading: ChatGPT-4o answered expert-validated true/false questions on vertical root fractures and tooth cracks with 86.1% accuracy, outperforming ChatGPT-3.5 and Gemini in a 5,400-response repeated test. gain_evidence: with ChatGPT-4o achieving the highest accuracy (86.1%), followed by ChatGPT-3.5 (85.6%) and Gemini (81.8%) problem_reading: The same models showed very low inter-model agreement and significant temporal instability, with Gemini performing worse in evening sessions, indicating they cannot replace clinician judgment for vertical root fracture decisions. problem_evidence: Inter-model agreement was very low (Cohen's kappa (κ) = 0.019) | evening performance was significantly lower for Gemini (p = 0.011) | limited agreement and temporal variability indicate that these systems cannot replace clinician judgment quick_read: A peer-reviewed study in Odontology compared ChatGPT-3.5, ChatGPT-4o, and Google Gemini on 60 expert-validated true/false questions about longitudinal tooth fractures, querying each model three times daily for 10 days for 5,400 total responses against expert reference answers. The results matter because even the best model reached only 86.1% accuracy with very low agreement across models and measurable instability over time, leaving uncertainty about reliability in real clinical workflows and reinforcing that such tools should remain supportive rather than determinative in endodontics. limitation: Findings are limited to dichotomous true/false questions derived from a single position statement, with very low inter-model agreement and significant temporal variability, supporting use only as supportive tools. tag: Dual reading key_points: Study used 60 validated true/false items based on European Society of Endodontology position statement on longitudinal tooth fractures, categorized into easy, moderate, and difficult levels. | Each of three models was queried over 10 consecutive days at three daily time points, generating 5,400 responses total. | Overall accuracy differed significantly among models (p < 0.001) and differences between models were evident only for moderate-difficulty questions (p < 0.001). | Inter-model agreement was very low with Cohen's kappa (k) = 0.019 and temporal stability differed significantly among models (p < 0.001). rundown: Researchers tested ChatGPT-3.5, ChatGPT-4o, and Google Gemini on 60 validated true/false questions drawn from the European Society of Endodontology position statement, with item-level content validity index of 1.00, administered over 10 days at morning, afternoon, and evening time points. Results showed ChatGPT-4o at 86.1% accuracy, ChatGPT-3.5 at 85.6%, and Gemini at 81.8%, with no morning-afternoon difference but significantly lower evening performance for Gemini, and very low inter-model agreement despite acceptable accuracy levels. sources: - peer_reviewed | Odontology | https://doi.org/10.1007/s10266-026-01570-6 | 2026-10-04 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- b2a6b71ce0f34b0913bc85364c262daf24cf7dff9ab2ec13b1fc8383b92be42c
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1305 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace