A Privacy‐Preserving and Explainable Deep Learning Framework for Imbalanced Student Performance Prediction Using CTGAN‐Based Data Augmentation

Citation

Tajdar, Samira and Dustmohammadloo, Hakimeh and Rakhmonov, Rauf and Mamadiyarov, Zokir and Urazbaeva, Dilbar and Tursunov, Lochin and Khamraev, Abror and Safoyev, Husen and Eldor, Tangirov (2026) A Privacy‐Preserving and Explainable Deep Learning Framework for Imbalanced Student Performance Prediction Using CTGAN‐Based Data Augmentation. Concurrency and Computation: Practice and Experience, 38 (15). ISSN 1532-0626

[img] Text
APRIVA~1.PDF - Published Version
Restricted to Repository staff only

Download (840kB)

Abstract

Identifying high-performing students in severely imbalanced educational datasets poses significant challenges for conventional machine learning models, which often optimize for overall accuracy while failing to detect extreme minorities. This study proposes a privacy-preserving and explainable deep learning framework that integrates Conditional Tabular Generative Adversarial Network (CTGAN)-based augmentation, adaptive model selection, and gradient-based interpretability for small-scale tabular data. Applied to a Student Physical Education Performance dataset of 500 records with only 2.4% High Performers, the framework synthesizes realistic minority-class samples via CTGAN and trains a heavily regularized Multi-Layer Perceptron (MLP) under strict early stopping. The proposed model was evaluated against six contemporary baselines: Logistic Regression, XGBoost, LightGBM, CatBoost, Soft Voting Ensemble, and TabTransformer under identical preprocessing, stratified splits, and equivalent hyperparameter budgets. A comprehensive imbalance-sensitive metric suite, including balanced accuracy, macro F1-score, ROC-AUC, and PR-AUC, was reported alongside per-class diagnostics. Systematic ablation studies comparing raw data, class weighting, SMOTE, and CTGAN confirm that generative augmentation is the critical driver of minority-class detection. Fivefold stratified cross-validation and computational efficiency analyses further validate robustness and the feasibility of deployment. The Final MLP achieved perfect High Performer recall (1.00), outperforming all baselines, while Integrated Gradients revealed that behavioral factors, particularly motivation and attendance, are stronger predictors than physical metrics. These findings establish a reproducible, transparent paradigm for imbalanced educational analytics.

Item Type: Article
Uncontrolled Keywords: Data augmentation
Subjects: Q Science > QA Mathematics > QA299.6-433 Analysis
Divisions: Faculty of Management (FOM)
Depositing User: Ms Rosnani Abd Wahab
Date Deposited: 02 Sep 2026 06:26
Last Modified: 02 Sep 2026 06:26
URII: http://shdl.mmu.edu.my/id/eprint/16529

Downloads

Downloads per month over past year

View ItemEdit (login required)