AI chatbots reporting jaw lesions from radiographic images
Source article: Accuracy of Artificial Intelligence based chatbots in reporting jaw lesions from multimodal radiographic images: A cross-sectional study
Objectives The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Materials and methods Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstruc…
Contested: both sides are scored from claims and sources, not community votes.

"Yuki the annoying chatbot" by danbri is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.
A cross-sectional study tested four AI chatbots on 97 anonymized CBCT cases of jaw lesions, comparing performance on reconstructed 2D panoramic views and, for Manus, raw 3D DICOM data. Reports were scored for accuracy, relevance and feasibility, revealing statistically significant differences between systems.
The findings matter because chatbot-generated radiology reports could influence clinical decisions for jaw pathology, yet performance varied sharply by model and input type. While 3D input improved results for one architecture, the study leaves open how these tools would perform prospectively, across broader populations, and under clinical safety oversight.
- Cross-sectional study used CBCT dataset from 97 patients with jaw lesions, anonymized and reconstructed into panoramic 2D views.
- Four chatbots tested: Gemini 2.5 Pro, Copilot, Claude, and Manus, with Manus also receiving raw DICOM data.
- Gemini 2.5 Pro ranked second with 80% of lesions detected and 56% correctly diagnosed.
- Reports were evaluated for accuracy, relevance and feasibility, with statistically significant differences across chatbots.
Manus architecture using raw 3D CBCT data detected and correctly diagnosed 95% of jaw lesions in 97 patients, outperforming 2D panoramic inputs.
Copilot and Claude produced the least accurate reports for jaw lesions, highlighting significant discrepancies in diagnostic accuracy across chatbots.
The rundown
Researchers collected anonymized CBCT data from 97 patients with jaw lesions and reconstructed panoramic 2D views using Bluesky Plan software. Those images were provided to Gemini 2.5 Pro, Copilot, Claude and Manus, while raw DICOM data was also provided to Manus with prompting.
Evaluation of generated reports found statistically significant differences in all measured parameters. Manus CBCT was most accurate, followed by Gemini 2.5 Pro, then Manus Pan, with Copilot and Claude lowest, supporting the conclusion that raw 3D data improves performance.
Sources
- Peer-reviewedDentomaxillofacial Radiology2026-08-04
How should this claim be treated?
ace
The debate