TRV-2026-0653Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0653 version: 1 kind: certified reason: Certified into the record timestamp: 2026-08-05T06:27:05.213466Z status: published lens: trace sector: health headline: Accuracy of Artificial Intelligence based chatbots in reporting jaw lesions from multimodal radiographic images: A cross-sectional study dek: Objectives The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Materials and methods Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstruc… gain_title: Manus architecture using raw 3D CBCT data detected and correctly diagnosed 95% of jaw lesions in 97 patients, outperforming 2D panoramic inputs. problem_title: Copilot and Claude produced the least accurate reports for jaw lesions, highlighting significant discrepancies in diagnostic accuracy across chatbots. trace_subject: AI chatbots reporting jaw lesions from radiographic images gain_reading: Manus architecture using raw 3D CBCT data detected and correctly diagnosed 95% of jaw lesions in 97 patients, outperforming 2D panoramic inputs. gain_evidence: 95% of lesions were detected and correctly diagnosed, | integration of raw 3-dimensional CBCT data substantially optimizes chatbot performance in lesion detection and diagnosis, as demonstrated by Manus architecture. problem_reading: Copilot and Claude produced the least accurate reports for jaw lesions, highlighting significant discrepancies in diagnostic accuracy across chatbots. problem_evidence: The least accurate results were detected in Copilot and Claude. | Significant discrepancies exist among artificial intelligence (AI) chatbots regarding their diagnostic accuracy in reporting jaw lesions. quick_read: A cross-sectional study tested four AI chatbots on 97 anonymized CBCT cases of jaw lesions, comparing performance on reconstructed 2D panoramic views and, for Manus, raw 3D DICOM data. Reports were scored for accuracy, relevance and feasibility, revealing statistically significant differences between systems. The findings matter because chatbot-generated radiology reports could influence clinical decisions for jaw pathology, yet performance varied sharply by model and input type. While 3D input improved results for one architecture, the study leaves open how these tools would perform prospectively, across broader populations, and under clinical safety oversight. limitation: tag: Automated dual reading key_points: Cross-sectional study used CBCT dataset from 97 patients with jaw lesions, anonymized and reconstructed into panoramic 2D views. | Four chatbots tested: Gemini 2.5 Pro, Copilot, Claude, and Manus, with Manus also receiving raw DICOM data. | Gemini 2.5 Pro ranked second with 80% of lesions detected and 56% correctly diagnosed. | Reports were evaluated for accuracy, relevance and feasibility, with statistically significant differences across chatbots. rundown: Researchers collected anonymized CBCT data from 97 patients with jaw lesions and reconstructed panoramic 2D views using Bluesky Plan software. Those images were provided to Gemini 2.5 Pro, Copilot, Claude and Manus, while raw DICOM data was also provided to Manus with prompting. Evaluation of generated reports found statistically significant differences in all measured parameters. Manus CBCT was most accurate, followed by Gemini 2.5 Pro, then Manus Pan, with Copilot and Claude lowest, supporting the conclusion that raw 3D data improves performance. sources: - peer_reviewed | Dentomaxillofacial Radiology | https://doi.org/10.1093/dmfr/twag057 | 2026-08-04 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- 3010787d0715991b16756504cd0abb10e410f9214a508742b17674873e296021
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0653 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace