Application of Multimodal Technology in Image Text Retrieval
摘要
In the field of artificial intelligence, multimodal technology is widely used in various fields such as image recognition, speech recognition, and natural language processing. This article first introduces the concept of multimodality, with a focus on the concept and implementation steps of graphic and textual retrieval tasks. It then introduces the principles and differences between CLIP and CNCLIP technologies, followed by how to implement graphic and textual retrieval through fine-tuning CNCLIP, including text retrieval of image and image retrieval of text. Finally, the advantages and disadvantages of CLIP are summarized. The experimental data in this article are all from the 2024 (12th) “Teddy Cup” national data mining challenge.