Deployment-Aware 30-Day Readmission Prediction in Resource-Limited Hospitals: Calibration, Threshold Policy, and Decision Utility

Citation

Malalha, Samer Asad and Burhanuddin, Ma and Duhair, Hatem T M and Alsayaydeh, Jamil Abedalrahim Jamil and Farid, Mazen (2026) Deployment-Aware 30-Day Readmission Prediction in Resource-Limited Hospitals: Calibration, Threshold Policy, and Decision Utility. International Journal of Advanced Computer Science and Applications, 17 (6). ISSN 2158-107X

[img] Text
2.pdf - Published Version
Restricted to Repository staff only

Download (963kB)

Abstract

Thirty-day hospital readmission is a well-established quality metric, and many clinical prediction models have been developed for this task; however, high discrimination does not by itself mean that a model is safe to use in discharge workflows. This study developed and applied an integrated deployment-oriented evaluation workflow in which calibration, threshold governance, and decision utility were treated as primary evaluation requirements rather than secondary diagnostics. Retrospective inpatient data collected between 2022 and 2024 from two resource-limited government hospitals were used (N = 30,000; readmission prevalence = 15.0%). The analysis was based on patient-level internal validation using non-overlapping training, validation, and held-out test partitions. A multilayer perceptron neural network and a random forest were evaluated using patient-level grouping (70% training, 15% validation, 15% test). Both models showed strong discrimination on the held-out test set (ROC-AUC = 0.868 for the neural network and 0.880 for the random forest), with nearly identical minority class detection (PR-AUC = 0.461 vs 0.460). However, calibration analyses separated the models despite identical Brier scores (0.095). The neural network showed lower expected calibration error (ECE = 0.013 vs 0.035) and near ideal probability scaling (calibration slope = 0.971, 95% CI: [0.94, 1.00]) compared with the random forest (slope = 1.202). Threshold analysis also showed that a default threshold could be unsafe, since recall was 0.244 at 0.50 but increased to 0.867 at 0.12, while false negatives dropped from 510 to 90. Decision Curve Analysis further supported the neural network, including a mean net benefit of 0.097 at a threshold of 0.12. Practically, threshold, model version, and monthly calibration summaries should be logged in an audit trail.

Item Type: Article
Uncontrolled Keywords: Clinical AI deployment, hospital readmission prediction, expected calibration error, electronic health records, operating region, probability reliability, model governance
Subjects: R Medicine > R Medicine (General) > R855-855.5 Medical technology
Divisions: Faculty of Information Science and Technology (FIST)
Depositing User: Ms Suzilawati Abu Samah
Date Deposited: 04 Aug 2026 01:53
Last Modified: 04 Aug 2026 01:53
URII: http://shdl.mmu.edu.my/id/eprint/16450

Downloads

Downloads per month over past year

View ItemEdit (login required)