Clinically Interpretable Deep Learning for Differentiating Vitiligo and Postinflammatory Hypopigmentation: Diagnostic Accuracy Study
Distinguishing vitiligo from postinflammatory hypopigmentation (PIH) is clinically challenging because both conditions may present with similar depigmented lesions. Although deep learning has shown strong potential for dermatologic image classification, limited interpretability remains a barrier to clinical adoption. This study aimed to develop an interpretable deep learning framework for accurate differentiation between vitiligo and PIH using a lightweight convolutional neural network and an ensemble of explain…
Hospital Universitari Doctor Josep Trueta de Girona by Premsaicsgirona. CC BY-SA 3.0 · https://creativecommons.org/licenses/by-sa/3.0
A diagnostic accuracy study published July 24, 2026 developed an interpretable deep learning framework to distinguish vitiligo from postinflammatory hypopigmentation, two conditions with similar depigmented lesions. Using 332 clinical images from King Abdullah University Hospital and public sources, a fine-tuned MobileNetV2 was evaluated with patient-wise 5-fold cross-validation.
By the publication date the model had demonstrated high measured performance and improved transparency through an ensemble of three explanation methods, suggesting potential support for dermatologists in differential diagnosis. What remains unproven is generalizability beyond the limited dataset and real-world clinical workflow impact.
- 332 clinical images (176 vitiligo and 156 PIH) collected from King Abdullah University Hospital and public online sources
- Patient-wise 5-fold cross-validation used to eliminate patient-level data leakage
- Pretrained MobileNetV2 fine-tuned by unfreezing final 30 layers
- Interpretability via equal-weight ensemble of Grad-CAM, integrated gradients, and SmoothGrad
A fine-tuned MobileNetV2 model differentiated vitiligo from postinflammatory hypopigmentation with 94.88% accuracy and 0.9885 AUC while providing clinically meaningful ensemble explanations.
The rundown
Researchers collected 332 images and trained a lightweight MobileNetV2 model with patient-wise 5-fold cross-validation, fine-tuning the final 30 layers and evaluating accuracy, precision, recall, F1-score and AUC.
To address interpretability barriers, they combined gradient-weighted class activation mapping (Grad-CAM), integrated gradients, and smooth gradients into an equal-weight ensemble, reporting clinically meaningful explanations in 98.48% of a 66-image validation subset.
Sources
- Peer-reviewedJMIR Medical Informatics2026-07-24
How should this claim be treated?
ace
The debate