Refine Outline First: Mask-Based Point-Image Model for Efficient Point Cloud Completion
摘要
In this paper, we explore an end-to-end framework called ROF (Refine Outline First) for cross-modal point cloud fusion and multi-scale reconstruction, designed to address the common challenge of incomplete point clouds in practical applications. As a key representation of 3D spatial data, point clouds are often sparse and incomplete due to factors such as occlusion, sensor resolution limitations, and data collection constraints. ROF addresses this by introducing an outline refinement module for recovering missing outline information and enhancing fine-grained details. Furthermore, we introduce a symmetric multimodal architecture that treats point cloud and image modalities equally, facilitating their joint completion and optimizing feature fusion. A masking strategy is applied during inference to maintain training-inference consistency and reduce computational cost, achieving a 27.11% FLOPs reduction. ShapeNet‑ViPC dataset demonstrate that ROF achieves competitive performance in joint point cloud and image completion tasks.