Citation
Ng, Jing Xiang and Lim, Kian Ming and Lim, Jia Min and Lim, Heng Siong and Ong, Thian Song and Lee, Chin Poo TiSI-ViT: Text-independent Speaker Identification with Vision Transformer. IAENG International Journal of Computer Science, 53 (9). pp. 3554-3567. ISSN 1819-656X|
Text
IJCS_53_9_12.pdf - Published Version Restricted to Repository staff only Download (3MB) |
Abstract
Traditional speaker identification frameworks have predominantly utilized hand-crafted acoustic features or convolutional neural networks (CNNs) for localized feature extraction. This research introduces TiSI-ViT, a novel Text-independent Speaker Identification framework that leverages the Vision Transformer (ViT) architecture to model global contextual dependencies within audio spectrograms. Unlike CNN-based models, TiSI-ViT utilizes a self-attention mechanism to capture long-range temporal-spectral relationships, which are essential for distinguishing subtle vocal variations between speakers. To overcome the challenges of training on limited datasets, a transfer learning strategy is employed by fine-tuning Vision Transformer pre-trained on ImageNet-21k. The proposed method standardizes audio via stereo-to-mono conversion and resampling to 44100Hz before transforming signals into fixed-size logarithmic spectrograms. Experimental evaluations on the in-house English100 dataset which comprising 14,000 utterances from 140 speakers demonstrate that TiSI-ViT achieves a state-of-the-art testing accuracy of 99.00. This performance significantly outperforms existing methods, including CNN-GraySpec and MVSI-Net, by 4.43 and 5.68 respectively, establishing the Vision Transformer as a robust architecture for text-independent speaker recognition. Quantitative evaluations, including embedding and statistical analyses, confirm superior feature compactness and reliability over conventional CNN and hybrid models. Furthermore, t-SNE visualizations and confusion matrix analysis provide qualitative evidence of enhanced discriminative capability and minimal inter-class error. © 2026, International Association of Engineers. All rights reserved.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Speaker Identification, Spectrogram, Text independent, Vision Transformer, Fine-tuning |
| Subjects: | Q Science > Q Science (General) |
| Divisions: | Faculty of Information Science and Technology (FIST) |
| Depositing User: | Ms Suzilawati Abu Samah |
| Date Deposited: | 01 Oct 2026 02:12 |
| Last Modified: | 01 Oct 2026 02:18 |
| URII: | http://shdl.mmu.edu.my/id/eprint/16751 |
Downloads
Downloads per month over past year
Edit (login required) |
