DistillCvT: Self-supervised distillation of Convolutional Vision Transformers for few-shot fine-grained classification

Citation

Lim, Jia Min and Lim, Kian Ming and Lim, Jit Yan and Lee, Chin Poo and Goh, Pey Yun (2026) DistillCvT: Self-supervised distillation of Convolutional Vision Transformers for few-shot fine-grained classification. Image and Vision Computing, 174. p. 106119. ISSN 0262-8856

[img] Text
1-s2.0-S026288562600226X-main.pdf - Published Version
Restricted to Repository staff only

Download (2MB)

Abstract

Few-shot fine-grained image classification remains a challenging task due to the subtle inter-class variations and the scarcity of labeled data. Existing few-shot fine-grained methods often struggle to generalize effectively under low-data regimes and fail to discriminate between visually similar samples. To address these challenges, we propose DistillCvT, a two-stage framework to enhance model generalization and feature discrimination through self-supervised learning and self-distillation within prototypical network structure. Convolutional Vision Transformer (CvT) is employed as the feature extractor, which integrates convolutional locality with transformer-based global modeling to capture fine-grained details while preserving global context. DistillCvT mitigates these issues through a two-stage strategy. In the first stage, the model is trained using a prototypical distance-based classification loss, where class probabilities are computed from the distances between query samples and class prototypes, complemented by standard data augmentation to mitigate data scarcity and learn discriminative class representations. In the second stage, an identical student model is trained with a self-supervised jigsaw task to enhance feature discrimination and robustness by aligning support samples with their augmented query counterparts. Simultaneously, the teacher model from stage one guides it through logit-based self-distillation to further improve generalization. Experiments on CUB-200-2011, Stanford Dogs, and Stanford Cars demonstrate that DistillCvT consistently improves performance and achieves competitive results compared with recent state-of-the-art few-shot fine-grained image classification methods.

Item Type: Article
Uncontrolled Keywords: Few-shot learning, Fine-grained image classification, Self-supervised learning, Prototypical network, Self-distillation
Subjects: Q Science > QA Mathematics > QA71-90 Instruments and machines
Divisions: Faculty of Information Science and Technology (FIST)
Depositing User: Ms Suzilawati Abu Samah
Date Deposited: 04 Aug 2026 02:44
Last Modified: 04 Aug 2026 02:44
URII: http://shdl.mmu.edu.my/id/eprint/16464

Downloads

Downloads per month over past year

View ItemEdit (login required)