<p>Building large-scale 3D datasets for pre-training a 3D shape understanding model is costly and requires considerable resources including expensive equipment, highly skilled engineers and relatively long time. To address this challenge, we propose an efficient approach named Point Contrastive Captioner (PointCoCa) which pre-trains the model using generated 3D datasets. Our work consists of two major components: (1) a data generation pipeline, which combines a text-conditional 3D shape generative model and a systematic prompt generation process; (2) a 3D shape understanding model, which is a multi-modal point cloud-text contrastive learning model pre-trained on the generated 3D dataset. Instead of aligning the 3D data with a pre-trained CLIP model, we pre-train PointCoCa from scratch for a more straightforward point cloud-text representation. PointCoCa achieves a state-of-the-art zero-shot classification accuracy of 65.7% on ScanObjectNN dataset and competitive performance on the ModelNet datasets which shows the effectiveness of the proposed approach. Additionally, based on the encoder-decoder architecture, PointCoCa can be directly applied to the point cloud captioning task without further modification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PointCoCa: Point contrastive captioner pre-trained with prompt-driven datasets generation enhances point cloud shape understanding

  • Fenglin Liu,
  • Tao Zhou,
  • Wei Wang,
  • Jianliang Ma,
  • Dong Li,
  • Wei Xi

摘要

Building large-scale 3D datasets for pre-training a 3D shape understanding model is costly and requires considerable resources including expensive equipment, highly skilled engineers and relatively long time. To address this challenge, we propose an efficient approach named Point Contrastive Captioner (PointCoCa) which pre-trains the model using generated 3D datasets. Our work consists of two major components: (1) a data generation pipeline, which combines a text-conditional 3D shape generative model and a systematic prompt generation process; (2) a 3D shape understanding model, which is a multi-modal point cloud-text contrastive learning model pre-trained on the generated 3D dataset. Instead of aligning the 3D data with a pre-trained CLIP model, we pre-train PointCoCa from scratch for a more straightforward point cloud-text representation. PointCoCa achieves a state-of-the-art zero-shot classification accuracy of 65.7% on ScanObjectNN dataset and competitive performance on the ModelNet datasets which shows the effectiveness of the proposed approach. Additionally, based on the encoder-decoder architecture, PointCoCa can be directly applied to the point cloud captioning task without further modification.