Human activity recognition under low-light conditions using frozen CLIP with lightweight adaptation

Citation

Khan, Md. Ashik and Miah, Abu Saleh Musa and Al Farid, Fahmid and Rahim, Md. Abdur and Abdul Karim, Hezerul (2026) Human activity recognition under low-light conditions using frozen CLIP with lightweight adaptation. Discover Artificial Intelligence, 6 (1). ISSN 2731-0809

[img] Text
s44163-026-02020-6.pdf - Published Version
Restricted to Repository staff only

Download (3MB)

Abstract

Human activity recognition (HAR) in low-light conditions remains challenging due to photometric domain shift, reduced contrast, altered colour statistics, and noise patterns absent from standard RGB pre-training data, causing models to degrade even after fine-tuning. To address this, we propose a lightweight adaptation framework leveraging frozen CLIP representations that ensures computational efficiency while delivering competitive performance across the three evaluated benchmarks. Our approach extracts embeddings from uniformly sampled frames without backbone updates, avoiding the computational overhead of full finetuning. By employing lightweight parameter-efficient adapters, we enable effective adaptation with minimal computational overhead. This design yields measurable accuracy improvements under low-light photometric shift. For instance, temporal adapters with only 88.8K parameters improve ARID v1.5 accuracy from 68.37 ± 0.49% to 74.38 ± 0.13%, an absolute gain of 6.01 percentage points. We further demonstrate the benefits of this approach in other settings: (1) linear probing achieves 87.26 ± 0.50% accuracy on UCF101 at 17.44 GFLOPs, within 6% of fine-tuned SlowOnly (93.16% at 27.3 GFLOPs) with 36% fewer computations, highlighting frozen CLIP’s efficiency as a baseline for standard datasets; and (2) zero-shot classification shows a performance gap of approximately 15–46 percentage points across datasets, underscoring the necessity of task-specific adaptation. All lightweight adapters operate at essentially the same cost as the frozen backbone since temporal adapters, LoRA, attention pooling, and prompt tuning each add under 0.2% additional FLOPs. These results establish frozen vision-language models with lightweight adaptation as a scalable and efficient framework for HAR under photometric domain shift, demonstrated on low-light infrared and standard RGB benchmarks, achieving meaningful gains with no backbone retraining required.

Item Type: Article
Uncontrolled Keywords: Video action recognition, Low-light video analysis, Vision-language models
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7800-8360 Electronics > TK7885-7895 Computer engineering. Computer hardware
Divisions: Faculty of Artificial Intelligence & Engineering (FAIE)
Depositing User: Ms Rosnani Abd Wahab
Date Deposited: 05 Oct 2026 01:02
Last Modified: 05 Oct 2026 01:02
URII: http://shdl.mmu.edu.my/id/eprint/16847

Downloads

Downloads per month over past year

View ItemEdit (login required)