TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-10

AI chatbot interpretation of radiological congenital anomaly and tumor cases

Source article: How Well Do AI Chatbots Understand Abnormal Anatomy: A Comparative Study Using Congenital Anomalies and Tumor Cases

Abstract: Artificial Intelligence (AI) chatbots are becoming an efficient option to understand medical data and also help with clinical reasoning. There has been a recent progression in research of large language models and their ability to be used in the healthcare sector, such as radiological image analysis, and diagnostic support. There is however, little evidence supporting their ability to accurately understand abnormal anatomical conditions, such as congenital anomalies and tumor related changes in anatomy. To asses…

TRV-2026-1047Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 70The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 70The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
How Well Do AI Chatbots Understand Abnormal Anatomy: A Comparative Study Using Congenital Anomalies and Tumor Cases

Conversation with a Sam Harris chatbot (51169748740) by Steve Jurvetson from Los Altos, USA. Public domain

The quick read

On September 9, 2026, a peer-reviewed comparative study reported testing ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases split between congenital anomalies and tumors. Each system received the same questions and images and was scored for diagnostic accuracy and explanatory completeness, with ChatGPT scoring 19 correct, Gemini 17, and Copilot 16.

The findings matter because accurate understanding of abnormal anatomy could support radiological assessment, yet the observed confusion between similar anomalies and reduced detail in complex tumors shows why outputs still need clinician oversight. With only 20 cases tested, it remains uncertain how performance would hold across diverse patient populations, imaging modalities, and real-world workflows.

Main points
  • Comparative cross-sectional study tested ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases: 10 congenital anomalies and 10 tumor cases.
  • Performance ranked ChatGPT 19/20 correct (95.0%), Gemini 17/20 (85.0%), Copilot 16/20 (80.0%) based on diagnostic accuracy and completeness.
  • Authors concluded chatbots showed high ability but responses should be checked by qualified healthcare professionals before clinical use.
Gain

In a 20-case test of congenital anomalies and tumors, ChatGPT, Gemini and Copilot achieved 80-95% diagnostic accuracy with detailed anatomical descriptions, suggesting potential as supplementary radiological diagnostic support.

Problem

Chatbots sometimes confused similar congenital anomalies and provided less detailed anatomical descriptions in complex tumor cases, requiring caution and verification by qualified professionals before clinical use.

The rundown

Researchers presented three chatbots with standard questions and images for 20 cases and scored responses on diagnostic accuracy, anatomical description, embryological or pathological explanation, and completeness.

ChatGPT provided the most detailed anatomical and pathological explanations, while Copilot and Gemini showed high correct rates but occasional confusion between similar congenital anomalies and thinner descriptions in complex tumor cases.

What this doesn’t fix

Study evaluated only 20 representative cases (10 congenital, 10 tumor) using standard questions and images, limiting generalizability to broader clinical practice and more complex presentations.

Sources

Reader signal

How should this claim be treated?

The debate