PointCoCa: Point contrastive captioner pre-trained with prompt-driven datasets generation enhances point cloud shape understanding
摘要
Building large-scale 3D datasets for pre-training a 3D shape understanding model is costly and requires considerable resources including expensive equipment, highly skilled engineers and relatively long time. To address this challenge, we propose an efficient approach named Point Contrastive Captioner (PointCoCa) which pre-trains the model using generated 3D datasets. Our work consists of two major components: (1) a data generation pipeline, which combines a text-conditional 3D shape generative model and a systematic prompt generation process; (2) a 3D shape understanding model, which is a multi-modal point cloud-text contrastive learning model pre-trained on the generated 3D dataset. Instead of aligning the 3D data with a pre-trained CLIP model, we pre-train PointCoCa from scratch for a more straightforward point cloud-text representation. PointCoCa achieves a state-of-the-art zero-shot classification accuracy of 65.7% on ScanObjectNN dataset and competitive performance on the ModelNet datasets which shows the effectiveness of the proposed approach. Additionally, based on the encoder-decoder architecture, PointCoCa can be directly applied to the point cloud captioning task without further modification.