Citation
Zain, SM and Mahamud, Eram and Assaduzzaman, Md and Fahad, Nafiz and Liew, Tze Hui (2026) Do deep chest X-ray classifiers explain where they look? A multi-architecture faithfulness benchmark with false-positive attribution audit on VinDr-CXR. Intelligence-Based Medicine, 15. p. 100431. ISSN 2666-5212|
Text
1-s2.0-S266652122600089X-main.pdf - Published Version Restricted to Repository staff only Download (8MB) |
Abstract
Background Hospital teams and external reviewers increasingly inspect saliency overlays during pre-deployment chest X-ray AI validation. A plausible heatmap can accompany an incorrect prediction, misleading validators who conflate visual highlighting with model correctness. Objective We propose and apply a structured pre-deployment attribution audit (null-controlled false-positive review and attentive-versus-blind false-negative categorization) on VinDr-CXR across three configurations (DenseNet-121, ConvNeXtV2-Tiny, Swin-B + LoRA) under Integrated Gradients. This study offers a structured pre-deployment warning protocol, not a deployable classifier and not a claim of clinical readiness, for detecting post-hoc saliency that appears plausible despite incorrect predictions. Methods VinDr-CXR is a public chest-radiograph dataset in which multiple radiologists independently annotated bounding boxes for 14 thoracic findings. On a 15,000-image VinDr-CXR training subset with 2-of-3 reader-consensus boxes, we generated 1528 Integrated Gradients maps and evaluated annotation overlap (precision-mass, recall top-50%, AUC-mIoU). Type B false positives were compared to an analytic null; false negatives classified as attentive or blind. Robustness checks included LIME, weight randomisation, and ablation. Results Across configurations, saliency–box overlap remained modest on correct predictions (maximum precision-mass 0.153). Type B wrong-class false positives concentrated IG mass above the analytic null (pooled Δ +9.8%, p = 7.4e-16), with overlays highlighting real abnormal regions despite an incorrect label. Attentive false negatives reached 51.4% of Swin-B + LoRA misses, including all Cardiomegaly cases. DenseNet-121 marginally failed the Adebayo weight-randomisation check (ρ = 0.512), warranting caution relative to ConvNeXtV2-Tiny (ρ = 0.263) and Swin-B + LoRA (ρ = −0.164). Conclusions Within this single public dataset using coarse consensus boxes and post-hoc saliency, saliency-box overlap should not be treated as evidence of diagnostic correctness during pre-deployment review. Because the top-predicted-class IG rule excluded three clinically important findings (Pneumothorax, Consolidation, Atelectasis) and because all analyses rest on one dataset without external or prospective clinical validation, results describe a pre-deployment *warning signal*, not deployment readiness. Null-controlled attribution audit may complement, not replace, conventional error analysis before clinical deployment.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Explainability, Faithfulness, Chest X-ray, Integrated gradients, Vision transformer, VinDr-CXR, False-positive analysis |
| Subjects: | R Medicine > R Medicine (General) > R855-855.5 Medical technology |
| Divisions: | Faculty of Information Science and Technology (FIST) |
| Depositing User: | Ms Suzilawati Abu Samah |
| Date Deposited: | 04 Aug 2026 03:48 |
| Last Modified: | 04 Aug 2026 03:48 |
| URII: | http://shdl.mmu.edu.my/id/eprint/16473 |
Downloads
Downloads per month over past year
Edit (login required) |
