TruaceTracing the truth around AIWednesday, August 5, 2026
Health·The Trace·Automated dual reading·Published 2026-08-05

AI chatbots reporting jaw lesions from radiographic images

Source article: Accuracy of Artificial Intelligence based chatbots in reporting jaw lesions from multimodal radiographic images: A cross-sectional study

Objectives The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Materials and methods Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstruc…

TRV-2026-0653Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 73The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 73The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Accuracy of Artificial Intelligence based chatbots in reporting jaw lesions from multimodal radiographic images: A cross-sectional study

"Yuki the annoying chatbot" by danbri is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.

The quick read

A cross-sectional study tested four AI chatbots on 97 anonymized CBCT cases of jaw lesions, comparing performance on reconstructed 2D panoramic views and, for Manus, raw 3D DICOM data. Reports were scored for accuracy, relevance and feasibility, revealing statistically significant differences between systems.

The findings matter because chatbot-generated radiology reports could influence clinical decisions for jaw pathology, yet performance varied sharply by model and input type. While 3D input improved results for one architecture, the study leaves open how these tools would perform prospectively, across broader populations, and under clinical safety oversight.

Main points
  • Cross-sectional study used CBCT dataset from 97 patients with jaw lesions, anonymized and reconstructed into panoramic 2D views.
  • Four chatbots tested: Gemini 2.5 Pro, Copilot, Claude, and Manus, with Manus also receiving raw DICOM data.
  • Gemini 2.5 Pro ranked second with 80% of lesions detected and 56% correctly diagnosed.
  • Reports were evaluated for accuracy, relevance and feasibility, with statistically significant differences across chatbots.
Gain

Manus architecture using raw 3D CBCT data detected and correctly diagnosed 95% of jaw lesions in 97 patients, outperforming 2D panoramic inputs.

Problem

Copilot and Claude produced the least accurate reports for jaw lesions, highlighting significant discrepancies in diagnostic accuracy across chatbots.

The rundown

Researchers collected anonymized CBCT data from 97 patients with jaw lesions and reconstructed panoramic 2D views using Bluesky Plan software. Those images were provided to Gemini 2.5 Pro, Copilot, Claude and Manus, while raw DICOM data was also provided to Manus with prompting.

Evaluation of generated reports found statistically significant differences in all measured parameters. Manus CBCT was most accurate, followed by Gemini 2.5 Pro, then Manus Pan, with Copilot and Claude lowest, supporting the conclusion that raw 3D data improves performance.

Sources

Reader signal

How should this claim be treated?

The debate