<p>Driven by the advancements in artificial intelligence technologies, image-to-point-cloud registration (I2P) have made significant strides. Nevertheless, the dimensional differences in the features of points cloud and image continue to pose considerable challenges to their development. The primary challenge is the inability to leverage the features of one modality to augment those of another, which complicates latent space feature alignment. To address this challenge, we propose an I2P method named TFCT-I2P. Initially, we introduce a Three-Stream Fusion Network (TFN), which integrates color information from images with structural information from point clouds, facilitating the alignment of features from both modalities. Subsequently, to effectively mitigate patch-level misalignments introduced by the inclusion of color information, we design a Color-Aware Transformer (CAT). Finally, we conduct extensive experiments on 7Scenes, RGB-D Scenes V2, ScanNet V2, and a self-collected dataset. The results demonstrate that TFCT-I2P surpasses state-of-the-art methods. Therefore, we believe that the proposed TFCT-I2P contributes to the advancement of I2P registration. The source code will be released at <a href="https://github.com/muyao99/TFCT-I2P">https://github.com/muyao99/TFCT-I2P</a> soon.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Three stream fusion network with color aware transformer for image to point cloud registration

  • Muyao Peng,
  • Pei An,
  • Zichen Wan,
  • You Yang,
  • Qiong Liu

摘要

Driven by the advancements in artificial intelligence technologies, image-to-point-cloud registration (I2P) have made significant strides. Nevertheless, the dimensional differences in the features of points cloud and image continue to pose considerable challenges to their development. The primary challenge is the inability to leverage the features of one modality to augment those of another, which complicates latent space feature alignment. To address this challenge, we propose an I2P method named TFCT-I2P. Initially, we introduce a Three-Stream Fusion Network (TFN), which integrates color information from images with structural information from point clouds, facilitating the alignment of features from both modalities. Subsequently, to effectively mitigate patch-level misalignments introduced by the inclusion of color information, we design a Color-Aware Transformer (CAT). Finally, we conduct extensive experiments on 7Scenes, RGB-D Scenes V2, ScanNet V2, and a self-collected dataset. The results demonstrate that TFCT-I2P surpasses state-of-the-art methods. Therefore, we believe that the proposed TFCT-I2P contributes to the advancement of I2P registration. The source code will be released at https://github.com/muyao99/TFCT-I2P soon.