TruaceTracing the truth around AISaturday, September 12, 2026
TRV-2026-1047Version 1 · Certified

Written 2026-09-10 06:05:42 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1047
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-09-10T06:05:42.244678Z
status: published
lens: trace
sector: health
headline: How Well Do AI Chatbots Understand Abnormal Anatomy: A Comparative Study Using Congenital Anomalies and Tumor Cases
dek: Artificial Intelligence (AI) chatbots are becoming an efficient option to understand medical data and also help with clinical reasoning. There has been a recent progression in research of large language models and their ability to be used in the healthcare sector, such as radiological image analysis, and diagnostic support. There is however, little evidence supporting their ability to accurately understand abnormal anatomical conditions, such as congenital anomalies and tumor related changes in anatomy. To asses…
gain_title: In a 20-case test of congenital anomalies and tumors, ChatGPT, Gemini and Copilot achieved 80-95% diagnostic accuracy with detailed anatomical descriptions, suggesting potential as supplementary radiological diagnostic support.
problem_title: Chatbots sometimes confused similar congenital anomalies and provided less detailed anatomical descriptions in complex tumor cases, requiring caution and verification by qualified professionals before clinical use.
trace_subject: AI chatbot interpretation of radiological congenital anomaly and tumor cases
gain_reading: In a 20-case test of congenital anomalies and tumors, ChatGPT, Gemini and Copilot achieved 80-95% diagnostic accuracy with detailed anatomical descriptions, suggesting potential as supplementary radiological diagnostic support.
gain_evidence: ChatGPT had the most accuracy with 19 correct responses (95.0%) | indicating their potential use as a supplementary tool in radiological assessment and radiological diagnostic support
problem_reading: Chatbots sometimes confused similar congenital anomalies and provided less detailed anatomical descriptions in complex tumor cases, requiring caution and verification by qualified professionals before clinical use.
problem_evidence: there are a number of ways in which the responses generated by AI should be interpreted with caution and checked by qualified healthcare professionals before it could be used in clinical settings | only limited confusion with similar congenital anomalies | sometimes mixed it up with similar congenital anomalies and less detailed with anatomical description
quick_read: On September 9, 2026, a peer-reviewed comparative study reported testing ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases split between congenital anomalies and tumors. Each system received the same questions and images and was scored for diagnostic accuracy and explanatory completeness, with ChatGPT scoring 19 correct, Gemini 17, and Copilot 16.

The findings matter because accurate understanding of abnormal anatomy could support radiological assessment, yet the observed confusion between similar anomalies and reduced detail in complex tumors shows why outputs still need clinician oversight. With only 20 cases tested, it remains uncertain how performance would hold across diverse patient populations, imaging modalities, and real-world workflows.
limitation: Study evaluated only 20 representative cases (10 congenital, 10 tumor) using standard questions and images, limiting generalizability to broader clinical practice and more complex presentations.
tag: Dual reading
key_points: Comparative cross-sectional study tested ChatGPT, Gemini, and Microsoft Copilot on 20 radiological cases: 10 congenital anomalies and 10 tumor cases. | Performance ranked ChatGPT 19/20 correct (95.0%), Gemini 17/20 (85.0%), Copilot 16/20 (80.0%) based on diagnostic accuracy and completeness. | Authors concluded chatbots showed high ability but responses should be checked by qualified healthcare professionals before clinical use.
rundown: Researchers presented three chatbots with standard questions and images for 20 cases and scored responses on diagnostic accuracy, anatomical description, embryological or pathological explanation, and completeness.

ChatGPT provided the most detailed anatomical and pathological explanations, while Copilot and Gemini showed high correct rates but occasional confusion between similar congenital anomalies and thinner descriptions in complex tumor cases.
sources:
- peer_reviewed | Clinical Anatomy | https://doi.org/10.1002/ca.70210 | 2026-09-09
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
049e092348a4b987ba51158a9bb7db2906376d0e9e024a8811d68188134411a4
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1047 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.