An interpretable multi-instance learning method for accurate differentiation of malignant and benign laryngeal lesions in laryngoscopy
Background Laryngeal cancer is a significant global health issue with high mortality, and early diagnosis is critical for survival. Developing accurate diagnostic models for laryngoscopy can reduce potential repeated biopsies and lessen the patient burden, representing an urgent clinical need. However, existing artificial intelligence models often function as black boxes and are trained on single, pre-selected images, which does not reflect the clinical workflow where multiple images are assessed. Methods We con…

In brief
Background Laryngeal cancer is a significant global health issue with high mortality, and early diagnosis is critical for survival. Developing accurate diagnostic models for laryngoscopy can reduce potential repeated biopsies and lessen the patient burden, representing an urgent clinical need.
However, existing artificial intelligence models often function as black boxes and are trained on single, pre-selected images, which does not reflect the clinical workflow where multiple images are assessed. Methods We conducted a retrospective study on 611 patients who underwent white light endoscopy (WLE) examinations.
Main points
- Background Laryngeal cancer is a significant global health issue with high mortality, and early diagnosis is critical for survival.
- Developing accurate diagnostic models for laryngoscopy can reduce potential repeated biopsies and lessen the patient burden, representing an urgent clinical need.
- However, existing artificial intelligence models often function as black boxes and are trained on single, pre-selected images, which does not reflect the clinical workflow where multiple images are assessed.
The gain
An interpretable multi-instance learning method for accurate differentiation of malignant and benign laryngeal lesions in laryngoscopy: Results IMIL-Net achieved the highest diagnostic performance, with a mean area under the curve (AUC) of 0.975 (95% CI 0.959-0.991), accuracy of 0.915 (95% CI 0.883-0.947), sensitivity of 0.876 (95% CI 0.803-0.949), and specificity of 0.945 (95% CI 0.910-0.980).
Sources
- Peer-reviewedEuropean Archives of Oto-Rhino-Laryngology2026-08-20
ace
The debate