Hybrid PointNet with Attention-Augmented T-Net for Robust 3D Point Cloud Classification

Citation

Batyha, Radwan M. and Owida, Hamza Abu and Mashagba, Hamza A. and Vasudevan, Asokan and Hunitie, Mohammad Faleh Ahmmad and Abd. Aziz, Azlan and Al-Mashagba, Lara A. and Mohammad, Suleiman Ibrahim (2026) Hybrid PointNet with Attention-Augmented T-Net for Robust 3D Point Cloud Classification. Applied Mathematics & Information Sciences, 20 (5). pp. 1229-1236. ISSN 19350090

[img] Text
j0j877n320n8zh.pdf - Published Version
Restricted to Repository staff only

Download (1MB)

Abstract

Point cloud classification is a foundational task in 3D scene understanding, with applications in autonomous driving, robotics, and medical image analysis. PointNet established a principled framework for directly consuming unordered 3D point sets via shared Multi-Layer Perceptrons (MLPs) and a Spatial Transformer Network (T-Net) with orthogonal regularization. Despite its efficiency, PointNet’s reliance on global max-pooling inside the T-Net discards potentially informative inter-point context, leading to suboptimal feature alignment under challenging conditions. In this work, we propose AT-PointNet (Attention-augmented T-Net PointNet), a hybrid architecture that replaces the global max-pooling aggregation inside both T-Net modules with a lightweight MultiHead Self-Attention (MHSA) mechanism. This modification enables the network to selectively weight point contributions based on global relational context before computing the transformation matrix, yielding richer and more discriminative feature representations while preserving PointNet’s permutation invariance. We evaluate AT-PointNet on the ModelNet10 benchmark, achieving an overall classification accuracy of 92.4%—a +2.1 percentage point improvement over the vanilla PointNet baseline (90.3%) under identical training conditions. Ablation studies confirm that the MHSA module in the input T-Net contributes the largest accuracy gain (+1.4%), and that four attention heads represent the optimal configuration. Comprehensive per-class analysis, ROC curves, confusion matrix decompositions, and attention visualization demonstrate the interpretability advantages of the proposed architecture.

Item Type: Article
Uncontrolled Keywords: Spatial transformer network, ModelNet10, 3D deep learning, attention mechanism
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7800-8360 Electronics > TK7885-7895 Computer engineering. Computer hardware
Divisions: Faculty of Engineering and Technology (FET)
Depositing User: Ms Rosnani Abd Wahab
Date Deposited: 05 Oct 2026 01:26
Last Modified: 05 Oct 2026 01:26
URII: http://shdl.mmu.edu.my/id/eprint/16852

Downloads

Downloads per month over past year

View ItemEdit (login required)