In the field of industrial design, efficiently retrieving and utilizing existing 3D assets has always been a critical challenge. With the rapid development of multimodal large model technology, it has become feasible to achieve unified representations of text, images, and 3D models at different levels. However, existing 3D asset databases often lack intelligent retrieval mechanisms when faced with large volumes of data, making it difficult to achieve precise matching through natural language or image inputs. To address this issue, this paper proposes a 3D model retrieval method based on multimodal large models, which can quickly locate target 3D models through natural language descriptions or image inputs. This method leverages multimodal large models for feature extraction, analysis, updating, and validation without requiring specialized training on specific data. Additionally, vector retrieval technology is employed as an auxiliary means to further enhance the efficiency and effectiveness of the retrieval process. Furthermore, this paper introduces a method for constructing a 3D model database that can automatically organize 3D assets and generate an efficiently query able database. We constructed a multimodal dataset comprising images, text, and 3D models and used the proposed method to establish a 3D asset database. The retrieval performance of 3D models was tested using both textual and image inputs. Experimental results demonstrate that the proposed method excels in 3D model retrieval tasks, achieving a Top-1 accuracy rate of 91.6% for image retrieval, thereby fully validating the effectiveness of our approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

3D Model Retrieval with Large Multimodal Models

  • Bo Wang,
  • Sicheng He,
  • Peng Liu,
  • Yun Li

摘要

In the field of industrial design, efficiently retrieving and utilizing existing 3D assets has always been a critical challenge. With the rapid development of multimodal large model technology, it has become feasible to achieve unified representations of text, images, and 3D models at different levels. However, existing 3D asset databases often lack intelligent retrieval mechanisms when faced with large volumes of data, making it difficult to achieve precise matching through natural language or image inputs. To address this issue, this paper proposes a 3D model retrieval method based on multimodal large models, which can quickly locate target 3D models through natural language descriptions or image inputs. This method leverages multimodal large models for feature extraction, analysis, updating, and validation without requiring specialized training on specific data. Additionally, vector retrieval technology is employed as an auxiliary means to further enhance the efficiency and effectiveness of the retrieval process. Furthermore, this paper introduces a method for constructing a 3D model database that can automatically organize 3D assets and generate an efficiently query able database. We constructed a multimodal dataset comprising images, text, and 3D models and used the proposed method to establish a 3D asset database. The retrieval performance of 3D models was tested using both textual and image inputs. Experimental results demonstrate that the proposed method excels in 3D model retrieval tasks, achieving a Top-1 accuracy rate of 91.6% for image retrieval, thereby fully validating the effectiveness of our approach.