Latency-Aware Hybrid Transformer–Capsule Network for Audio-Visual Emotion Recognition in Edge–Fog–Cloud Environments

Citation

Shukla, Abhinav and Pahuja, Deepika and Agrawal, Ayush Kumar and Ramasamy, R Kanesaraj and Dubey, Parul (2026) Latency-Aware Hybrid Transformer–Capsule Network for Audio-Visual Emotion Recognition in Edge–Fog–Cloud Environments. Algorithms, 19 (8). p. 626. ISSN 1999-4893

[img] Text
algorithms-19-00626.pdf - Published Version
Restricted to Repository staff only

Download (4MB)

Abstract

Audio-visual emotion recognition (AVER) is central to affective computing systems that require reliable, real-time interpretation of human emotions. However, many existing multimodal models treat feature learning and deployment efficiency separately, limiting their ability to preserve hierarchical facial relationships, capture long-range speech dynamics, and operate with low latency in distributed settings. This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment. The visual stream employs a CNN–Capsule branch to retain spatial hierarchies in facial expressions, while the audio stream uses a CNN–Transformer branch to learn local spectral patterns and long-range temporal dependencies from speech. A cross-modal Transformer fusion module integrates complementary emotional cues, and a latency-aware task-allocation mechanism allocates preprocessing, inference, and training-related operations across edge, fog, and cloud layers according to workload, node capacity, and communication delay. Unlike approaches that optimize multimodal representation learning and distributed deployment as separate problems, the proposed framework adopts a deployment-aware co-design in which spatial visual representation, temporal acoustic modeling, multimodal interaction, and deterministic latency-aware task allocation are coordinated within a unified processing pipeline. The framework is evaluated on RAVDESS, CREMA-D, and SAVEE using a subject-independent protocol. Experimental results show an average accuracy of 91.5%, an F1-score of 90.7%, an MCC of 0.894, and an AUC of 0.950. The framework further incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources. Physical-device deployment and comprehensive resource profiling remain subjects for future validation.

Item Type: Article
Uncontrolled Keywords: audio-visual emotion recognition, transformer networks, capsule networks, multimodal deep learning, affective computing
Subjects: Q Science > QA Mathematics > QA71-90 Instruments and machines
Divisions: Faculty of Computing and Informatics (FCI)
Depositing User: Ms Suzilawati Abu Samah
Date Deposited: 03 Sep 2026 04:00
Last Modified: 03 Sep 2026 04:00
URII: http://shdl.mmu.edu.my/id/eprint/16603

Downloads

Downloads per month over past year

View ItemEdit (login required)