Citation
Khan, Md. Ashik and Miah, Abu Saleh Musa and Al Farid, Fahmid and Rahim, Md. Abdur and Abdul Karim, Hezerul (2026) Human activity recognition under low-light conditions using frozen CLIP with lightweight adaptation. Discover Artificial Intelligence, 6 (1). ISSN 2731-0809|
Text
s44163-026-02020-6.pdf - Published Version Restricted to Repository staff only Download (3MB) |
Abstract
Human activity recognition (HAR) in low-light conditions remains challenging due to photometric domain shift, reduced contrast, altered colour statistics, and noise patterns absent from standard RGB pre-training data, causing models to degrade even after fine-tuning. To address this, we propose a lightweight adaptation framework leveraging frozen CLIP representations that ensures computational efficiency while delivering competitive performance across the three evaluated benchmarks. Our approach extracts embeddings from uniformly sampled frames without backbone updates, avoiding the computational overhead of full finetuning. By employing lightweight parameter-efficient adapters, we enable effective adaptation with minimal computational overhead. This design yields measurable accuracy improvements under low-light photometric shift. For instance, temporal adapters with only 88.8K parameters improve ARID v1.5 accuracy from 68.37 ± 0.49% to 74.38 ± 0.13%, an absolute gain of 6.01 percentage points. We further demonstrate the benefits of this approach in other settings: (1) linear probing achieves 87.26 ± 0.50% accuracy on UCF101 at 17.44 GFLOPs, within 6% of fine-tuned SlowOnly (93.16% at 27.3 GFLOPs) with 36% fewer computations, highlighting frozen CLIP’s efficiency as a baseline for standard datasets; and (2) zero-shot classification shows a performance gap of approximately 15–46 percentage points across datasets, underscoring the necessity of task-specific adaptation. All lightweight adapters operate at essentially the same cost as the frozen backbone since temporal adapters, LoRA, attention pooling, and prompt tuning each add under 0.2% additional FLOPs. These results establish frozen vision-language models with lightweight adaptation as a scalable and efficient framework for HAR under photometric domain shift, demonstrated on low-light infrared and standard RGB benchmarks, achieving meaningful gains with no backbone retraining required.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Video action recognition, Low-light video analysis, Vision-language models |
| Subjects: | T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7800-8360 Electronics > TK7885-7895 Computer engineering. Computer hardware |
| Divisions: | Faculty of Artificial Intelligence & Engineering (FAIE) |
| Depositing User: | Ms Rosnani Abd Wahab |
| Date Deposited: | 05 Oct 2026 01:02 |
| Last Modified: | 05 Oct 2026 01:02 |
| URII: | http://shdl.mmu.edu.my/id/eprint/16847 |
Downloads
Downloads per month over past year
Edit (login required) |
