错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cascaded Bilinear Mapping Collaborative Hybrid Attention Modality Fusion Model

  • Jiayu Zhao,
  • Kuizhi Mei

摘要

LiDAR and Camera feature fusion detection is widely used in the field of 3D object detection. A fusion method that integrates the two modalities into a unified representation has been proposed and proven effective. However, there is still a challenge in fusing these two modalities in the same space. To address this issue, we propose a fusion module that leverages bilinear mapping technology and a hybrid attention mechanism. Our method preserves the original independent branches for LiDAR and Camera feature extraction. We then employ a cascade bilinear mapping hybrid attention fusion module to combine the features of these two modalities in the Bird’s Eye View (BEV) space. We utilize a hybrid structure of Convolutional Neural Networks (CNN) and Efficient-Attention to process the fused features. Furthermore, Multilayer Perceptron (MLP) is used to extract information from the features and achieve an asymmetric deep fusion of the two modalities in BEV space through the cascaded bilinear mapping method. This approach helps to complement the information provided by both features and achieve improved detection results. On the nuScenes validation set, our model improves the accuracy by 1.2 \(\%\) compared to the current Sota method of feature fusion in BEV space. It is important to note that this improvement is achieved without any data augmentation or Test-Time Augmentation (TTA) techniques.