TRV-2026-1265Version 1 · Certified

Written 2026-10-03 06:56:35 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1265
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-10-03T06:56:35.118451Z
status: published
lens: p_space
sector: health
headline: Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal
dek: Objectives To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 checklist and STARD-Artificial Intelligence (AI)/Developmental and Exploratory Clinical Investigations of…
gain_title: (none)
problem_title: Despite preserved discrimination, the algorithm showed calibration failure with compressed scores, low positive predictive value and low agreement, indicating distributional shift in this low- and middle-income country setting.
trace_subject: (none)
gain_reading: (none)
gain_evidence: (none)
problem_reading: Despite preserved discrimination, the algorithm showed calibration failure with compressed scores, low positive predictive value and low agreement, indicating distributional shift in this low- and middle-income country setting.
problem_evidence: positive predictive value 17.89% (95% Wilson CI 13.4% to 23.5%) and Cohen's κ 0.237 | Systematic score compression-preserved discrimination despite calibration shift-is a quantifiable marker of low- and middle-income country distributional shift.
quick_read: From 5 to 20 June 2026, researchers prospectively tested a publicly available DenseNet-121 model from TorchXRayVision in shadow mode on 826 consecutive chest radiographs from health assessment applicants at Patan Hospital in Lalitpur, Nepal, comparing AI scores to blinded single-reader radiologist classifications.

The model preserved discrimination with high sensitivity and NPV, suggesting rule-out triage potential, but exhibited systematic calibration failure and low PPV in this population, raising questions about threshold stability, reference standard robustness, and readiness for operational deployment without local recalibration and independent external validation.
limitation: Single-reader reference standard, small number of positives creating wide uncertainty, and lack of independent external validation limit generalizability before operational deployment.
tag: Evidence-backed problem
key_points: Prospective single-centre shadow-mode study from 5 June 2026 to 20 June 2026 inclusive with 826 consecutive applicants; 41 (4.97%) classified abnormal by reference standard. | Index test was TorchXRayVision densenet121-res224-all with maximum aggregated pathology probability compared against post hoc threshold 0.6258 selected for >=95% sensitivity. | Reference standard was single-reader-per-case review by two board-certified radiodiagnosticians and one resident, blinded to AI output. | 10-fold cross-validation showed bias-corrected specificity 75.80% with optimism +1.40 pp; independent external validation was not performed.
rundown: The study enrolled 826 consecutive foreign employment predeparture and student migration applicants at Patan Academy of Health Sciences/Patan Hospital, with two exclusions for DICOM technical failure, and used a maximum aggregated pathology probability score.

At threshold 0.6258, specificity was 77.2% (95% CI 74.1% to 80.0%) with PPV 17.89% and Cohen's kappa 0.237, while Brier score 0.3621 exceeded null Brier 0.0472 and ECE was 0.564, showing score compression from 0.52-0.72.

Authors concluded high sensitivity supports potential as radiographic abnormality rule-out triage tool but noted this does not constitute microbiological exclusion of active pulmonary tuberculosis and called for prospective local calibration and external validation.
sources:
- peer_reviewed | BMJ Open | https://doi.org/10.1136/bmjopen-2026-124868 | 2026-10-01
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
749b5f04d9603132da0ab2d3e5d55b8957476c704812a8c3eaa6e876e5e3e3f3
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1265 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.