Efficient Point Transformer with Dynamic Token Aggregating for LiDAR Point Cloud Processing

Lu, Dening; Zhou, Jun; Kyle; Gao; Xu, Linlin; Li, Jonathan

Computer Science > Computer Vision and Pattern Recognition

arXiv:2405.15827 (cs)

[Submitted on 23 May 2024 (v1), last revised 25 Dec 2024 (this version, v2)]

Title:Efficient Point Transformer with Dynamic Token Aggregating for LiDAR Point Cloud Processing

Authors:Dening Lu, Jun Zhou, Kyle (Yilin)Gao, Linlin Xu, Jonathan Li

View PDF

Abstract:Recently, LiDAR point cloud processing and analysis have made great progress due to the development of 3D Transformers. However, existing 3D Transformer methods usually are computationally expensive and inefficient due to their huge and redundant attention maps. They also tend to be slow due to requiring time-consuming point cloud sampling and grouping processes. To address these issues, we propose an efficient point TransFormer with Dynamic Token Aggregating (DTA-Former) for point cloud representation and processing. Firstly, we propose an efficient Learnable Token Sparsification (LTS) block, which considers both local and global semantic information for the adaptive selection of key tokens. Secondly, to achieve the feature aggregation for sparsified tokens, we present the first Dynamic Token Aggregating (DTA) block in the 3D Transformer paradigm, providing our model with strong aggregated features while preventing information loss. After that, a dual-attention Transformer-based Global Feature Enhancement (GFE) block is used to improve the representation capability of the model. Equipped with LTS, DTA, and GFE blocks, DTA-Former achieves excellent classification results via hierarchical feature learning. Lastly, a novel Iterative Token Reconstruction (ITR) block is introduced for dense prediction whereby the semantic features of tokens and their semantic relationships are gradually optimized during iterative reconstruction. Based on ITR, we propose a new W-net architecture, which is more suitable for Transformer-based feature learning than the common U-net design.

Comments:	16 pages, 12 figures, 7 tables
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2405.15827 [cs.CV]
	(or arXiv:2405.15827v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2405.15827

Submission history

From: Dening Lu [view email]
[v1] Thu, 23 May 2024 20:50:50 UTC (7,114 KB)
[v2] Wed, 25 Dec 2024 06:20:14 UTC (7,786 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Efficient Point Transformer with Dynamic Token Aggregating for LiDAR Point Cloud Processing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.

Computer Science > Computer Vision and Pattern Recognition

Title:Efficient Point Transformer with Dynamic Token Aggregating for LiDAR Point Cloud Processing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.