TiSI-ViT: Text-independent Speaker Identification with Vision Transformer

Citation

Ng, Jing Xiang and Lim, Kian Ming and Lim, Jia Min and Lim, Heng Siong and Ong, Thian Song and Lee, Chin Poo TiSI-ViT: Text-independent Speaker Identification with Vision Transformer. IAENG International Journal of Computer Science, 53 (9). pp. 3554-3567. ISSN 1819-656X

[img] Text
IJCS_53_9_12.pdf - Published Version
Restricted to Repository staff only

Download (3MB)

Abstract

Traditional speaker identification frameworks have predominantly utilized hand-crafted acoustic features or convolutional neural networks (CNNs) for localized feature extraction. This research introduces TiSI-ViT, a novel Text-independent Speaker Identification framework that leverages the Vision Transformer (ViT) architecture to model global contextual dependencies within audio spectrograms. Unlike CNN-based models, TiSI-ViT utilizes a self-attention mechanism to capture long-range temporal-spectral relationships, which are essential for distinguishing subtle vocal variations between speakers. To overcome the challenges of training on limited datasets, a transfer learning strategy is employed by fine-tuning Vision Transformer pre-trained on ImageNet-21k. The proposed method standardizes audio via stereo-to-mono conversion and resampling to 44100Hz before transforming signals into fixed-size logarithmic spectrograms. Experimental evaluations on the in-house English100 dataset which comprising 14,000 utterances from 140 speakers demonstrate that TiSI-ViT achieves a state-of-the-art testing accuracy of 99.00. This performance significantly outperforms existing methods, including CNN-GraySpec and MVSI-Net, by 4.43 and 5.68 respectively, establishing the Vision Transformer as a robust architecture for text-independent speaker recognition. Quantitative evaluations, including embedding and statistical analyses, confirm superior feature compactness and reliability over conventional CNN and hybrid models. Furthermore, t-SNE visualizations and confusion matrix analysis provide qualitative evidence of enhanced discriminative capability and minimal inter-class error. © 2026, International Association of Engineers. All rights reserved.

Item Type: Article
Uncontrolled Keywords: Speaker Identification, Spectrogram, Text independent, Vision Transformer, Fine-tuning
Subjects: Q Science > Q Science (General)
Divisions: Faculty of Information Science and Technology (FIST)
Depositing User: Ms Suzilawati Abu Samah
Date Deposited: 01 Oct 2026 02:12
Last Modified: 01 Oct 2026 02:18
URII: http://shdl.mmu.edu.my/id/eprint/16751

Downloads

Downloads per month over past year

View ItemEdit (login required)