Point cloud completion aims to fill in incomplete or partially missing point cloud data to restore its complete shape information. Currently, some methods attempt to incorporate image modality information to achieve high-quality point cloud completion. However, they fail to fully integrate complementary features within the multimodal context. In this paper, we propose a novel Residual Multimodal Fusion network for point cloud completion, significantly improving the quality of shape completion. Specifically, we introduce a cross-modal residual feature fusion module to capture local shape features. It uses cross-modal attention mechanisms, while employing a residual structure to mitigate the process of globalizing features, thus effectively enhancing feature diversity. The decoder adopts an innovative attention-based multi-branch structure to reconstruct the complete point cloud by regions. Additionally, the point cloud refinement module is divided into local refinement units and view-assisted units, which can simultaneously capture global shape structures and local details, reducing outliers in the predicted point cloud. Experiments show that our network achieves competitive performance on synthetic and real-world datasets, outperforming existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Point Cloud Completion via Residual Attention Feature Fusion

  • Junkang Wan,
  • Hang Wu,
  • Yubin Miao

摘要

Point cloud completion aims to fill in incomplete or partially missing point cloud data to restore its complete shape information. Currently, some methods attempt to incorporate image modality information to achieve high-quality point cloud completion. However, they fail to fully integrate complementary features within the multimodal context. In this paper, we propose a novel Residual Multimodal Fusion network for point cloud completion, significantly improving the quality of shape completion. Specifically, we introduce a cross-modal residual feature fusion module to capture local shape features. It uses cross-modal attention mechanisms, while employing a residual structure to mitigate the process of globalizing features, thus effectively enhancing feature diversity. The decoder adopts an innovative attention-based multi-branch structure to reconstruct the complete point cloud by regions. Additionally, the point cloud refinement module is divided into local refinement units and view-assisted units, which can simultaneously capture global shape structures and local details, reducing outliers in the predicted point cloud. Experiments show that our network achieves competitive performance on synthetic and real-world datasets, outperforming existing methods.