Citation
Abuowaida, Suhaila and Owida, Hamza Abu and Mashagba, Hamza A. and Abd. Aziz, Azlan and Alali, Muath and Mashagba, Mohammadnour (2026) Deep Transfer Learning with Two-Stage Fine-Tuning for Robust Concrete Crack Detection. ES Energy & Environment. ISSN 25780646|
Text
Deep Transfer Learning with Two-Stage Fine-Tuning for Robust Concrete Crack Detection.pdf - Published Version Restricted to Repository staff only Download (1MB) |
Abstract
This work addresses automated crack detection in concrete structures as part of structural health monitoring (SHM) and introduces a benchmark based on recent advancements in transformer-based architectures as opposed to traditional convolutional neural networks (CNN). Recent CNN-based architectures include ResNet50, InceptionV3 and Xception, which while providing high quality results have been shown to be limited in their ability to capture spatial relationships over large areas, in their ability to classify fine cracks under changing illumination conditions, and in their ability to generalize to different surfaces of concrete structures. This research presents a transfer learning architecture where these previous backbones are replaced by three state-of-the-art transformer-based architectures: The Swin Transformer, EfficientNetV2-S, and ConvNeXtTiny. These models were each trained on the METU concrete crack dataset and the SDNET2018 multi-surface dataset, and both binary and multi-class classification problems were addressed. The datasets were preprocessed using a variety of methods to help address class imbalances present in the SDNET2018 dataset, including randomly sampling more images from the minority classes than from the majority classes. Results show that ConvNeXt-Tiny outperforms all previously published deep learning solutions on both the METU and SDNET2018 pavement subsets, achieving 99.98% accuracy on METU and 99.61% accuracy on the SDNET2018 pavement subset. In addition, the results demonstrate that EfficientNetV2-S provides the best trade-off of accuracy and computation, providing an inference time of 8.7 milliseconds per image at 99.94% accuracy on METU. Lastly, the results indicate that the Swin Transformer achieves competitive performance due to the hierarchical nature of its window-shifted self-attention mechanism, achieving 99.91% accuracy on METU and 98.14% accuracy on the full six-class SDNET2018 task. The contributions made by this study include (1) the first use of a two stage fine tuning technique to train Swin Transformer, EfficientNetV2-S and ConvNeXt-Tiny for the purposes of detecting cracks in concrete using images; (2) an experimentally verified theoretical method that is used to balance class representation within multi surface classification tasks where there are six classes; and (3), as compared to previous studies , new state of the art results were obtained when performing image based classification to detect cracks in concrete on both benchmark datasets.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Swin transformer, EfficientNetV2, ConvNeXt, Transfer learning |
| Subjects: | Q Science > QA Mathematics > QA71-90 Instruments and machines > QA75.5-76.95 Electronic computers. Computer science |
| Divisions: | Faculty of Engineering and Technology (FET) |
| Depositing User: | Ms Rosnani Abd Wahab |
| Date Deposited: | 03 Sep 2026 04:46 |
| Last Modified: | 03 Sep 2026 04:46 |
| URII: | http://shdl.mmu.edu.my/id/eprint/16610 |
Downloads
Downloads per month over past year
Edit (login required) |
