In the field of artificial intelligence, multimodal technology is widely used in various fields such as image recognition, speech recognition, and natural language processing. This article first introduces the concept of multimodality, with a focus on the concept and implementation steps of graphic and textual retrieval tasks. It then introduces the principles and differences between CLIP and CNCLIP technologies, followed by how to implement graphic and textual retrieval through fine-tuning CNCLIP, including text retrieval of image and image retrieval of text. Finally, the advantages and disadvantages of CLIP are summarized. The experimental data in this article are all from the 2024 (12th) “Teddy Cup” national data mining challenge.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Multimodal Technology in Image Text Retrieval

  • Yuyuan He,
  • Lihua Xiong,
  • Shujian Wang

摘要

In the field of artificial intelligence, multimodal technology is widely used in various fields such as image recognition, speech recognition, and natural language processing. This article first introduces the concept of multimodality, with a focus on the concept and implementation steps of graphic and textual retrieval tasks. It then introduces the principles and differences between CLIP and CNCLIP technologies, followed by how to implement graphic and textual retrieval through fine-tuning CNCLIP, including text retrieval of image and image retrieval of text. Finally, the advantages and disadvantages of CLIP are summarized. The experimental data in this article are all from the 2024 (12th) “Teddy Cup” national data mining challenge.