Transformer-Based Deep Learning Framework for Automated Lesion Detection in Capsule Endoscopy: A Comparative Study With CNN Architectures
Goals To compare a vision transformer with 2 convolutional neural network architectures for multiclass lesion classification in capsule endoscopy images. Background Manual review of capsule endoscopy is time-consuming and subject to interobserver variability. Deep learning can automate lesion recognition; however, most prior capsule endoscopy work evaluates a small number of classes, and systematic comparisons between transformer and convolutional architectures across many lesion categories are limited. Study Tw…
A pretrained Vision Transformer fine-tuned on a merged 21-class capsule endoscopy dataset achieved 92.2% accuracy and 0.99 AUC on an independent test set, outperforming DenseNet121 and ResNet50.
Because the dataset was split at the image level, correlated frames from the same examination may inflate performance, so results cannot be interpreted as patient-level generalization and require grouped reanalysis and external validation before clinical use.
Performance estimates are based on an image-level split rather than patient or procedure-level split, so frames from the same examination may be correlated and results should not be interpreted as patient-level generalization or definitive architectural superiority without grouped reanalysis and external validation.
Evidence
- Peer-reviewedJournal of Clinical Gastroenterology2026-08-14
How should this claim be treated?
Truvace Impact Record TRV-2026-0779, v1: “Transformer-Based Deep Learning Framework for Automated Lesion Detection in Capsule Endoscopy: A Comparative Study With CNN Architectures.” Truvace, 2026-08-16. /record/TRV-2026-0779 (accessed at citation time). sha256 8f1ed7e35a121157…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0779 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace