EndoVLM: A Vision-Language Assistant for Gastrointestinal Endoscopy
Gastrointestinal endoscopy generates extensive high-resolution video data, posing significant challenges for efficient and accurate computer-aided diagnosis of gastrointestinal diseases. To address this, we propose EndoVLM (Endoscopy Vision-Language Model), a specialized visual question-answering assistant for gastroenterology. EndoVLM introduces ConvNeXt as a hierarchical visual encoder to replace traditional ViTs (Vision Transformers), inherently compressing high-resolution gastrointestinal endoscopy images in…
EndoVLM increased visual question-answering accuracy for gastrointestinal endoscopy, raising Kvasir-VQA scores and improving external validation on Gastrovision.
Evidence
- Peer-reviewedJournal of Imaging Informatics in Medicine2026-08-26
How should this claim be treated?
Truvace Impact Record TRV-2026-0916, v1: “EndoVLM: A Vision-Language Assistant for Gastrointestinal Endoscopy.” Truvace, 2026-08-28. /record/TRV-2026-0916 (accessed at citation time). sha256 ae04103bbab9a737…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0916 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace