Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal
Objectives To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 checklist and STARD-Artificial Intelligence (AI)/Developmental and Exploratory Clinical Investigations of…

In brief
From 5 to 20 June 2026, researchers prospectively tested a publicly available DenseNet-121 model from TorchXRayVision in shadow mode on 826 consecutive chest radiographs from health assessment applicants at Patan Hospital in Lalitpur, Nepal, comparing AI scores to blinded single-reader radiologist classifications.
The model preserved discrimination with high sensitivity and NPV, suggesting rule-out triage potential, but exhibited systematic calibration failure and low PPV in this population, raising questions about threshold stability, reference standard robustness, and readiness for operational deployment without local recalibration and independent external validation.
Main points
- Prospective single-centre shadow-mode study from 5 June 2026 to 20 June 2026 inclusive with 826 consecutive applicants; 41 (4.97%) classified abnormal by reference standard.
- Index test was TorchXRayVision densenet121-res224-all with maximum aggregated pathology probability compared against post hoc threshold 0.6258 selected for >=95% sensitivity.
- Reference standard was single-reader-per-case review by two board-certified radiodiagnosticians and one resident, blinded to AI output.
- 10-fold cross-validation showed bias-corrected specificity 75.80% with optimism +1.40 pp; independent external validation was not performed.
The problem
Despite preserved discrimination, the algorithm showed calibration failure with compressed scores, low positive predictive value and low agreement, indicating distributional shift in this low- and middle-income country setting.
The rundown
The study enrolled 826 consecutive foreign employment predeparture and student migration applicants at Patan Academy of Health Sciences/Patan Hospital, with two exclusions for DICOM technical failure, and used a maximum aggregated pathology probability score.
At threshold 0.6258, specificity was 77.2% (95% CI 74.1% to 80.0%) with PPV 17.89% and Cohen's kappa 0.237, while Brier score 0.3621 exceeded null Brier 0.0472 and ECE was 0.564, showing score compression from 0.52-0.72.
Authors concluded high sensitivity supports potential as radiographic abnormality rule-out triage tool but noted this does not constitute microbiological exclusion of active pulmonary tuberculosis and called for prospective local calibration and external validation.
Sources
- Peer-reviewedBMJ Open2026-10-01
ace
The debate