Unsupervised methods have recently garnered significant attention owing to their effective retrieval capabilities and minimal storage requirements. However, current unsupervised techniques do not sufficiently capture the information about the joint occurrence of data and primarily rely on high-dimensional features to construct instance similarity matrices. This approach fails to effectively guide the learning of hash codes. To tackle the previously mentioned challenges, we put forward CLIP-Based Clustering Method for Unsupervised Hashing Multi-Modal Retrieval(CCUH). First, we extract pre-existing image and text features from the CLIP model to serve as the input data for the neural network. Then, we design a novel multimodal association matrix generator that employs cosine similarity and KMeans clustering algorithms. This generator combines the high-dimensional features obtained by the network with discriminative embeddings to build a cross-channel semantic association matrix enriched with additional information. This approach effectively supervises the learning of hash codes and improves the arrangement of semantic information across various modalities. Extensive empirical findings Indicate that on two commonly utilized datasets, our proposed CCUH method surpasses existing leading methods in cross-modal search tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CCUH: CLIP-Based Clustering Method for Unsupervised Hashing Multi-modal Retrieval

  • Qinze Zhu,
  • Xinsheng Shu,
  • Jiayi Yang,
  • Mingyong Li

摘要

Unsupervised methods have recently garnered significant attention owing to their effective retrieval capabilities and minimal storage requirements. However, current unsupervised techniques do not sufficiently capture the information about the joint occurrence of data and primarily rely on high-dimensional features to construct instance similarity matrices. This approach fails to effectively guide the learning of hash codes. To tackle the previously mentioned challenges, we put forward CLIP-Based Clustering Method for Unsupervised Hashing Multi-Modal Retrieval(CCUH). First, we extract pre-existing image and text features from the CLIP model to serve as the input data for the neural network. Then, we design a novel multimodal association matrix generator that employs cosine similarity and KMeans clustering algorithms. This generator combines the high-dimensional features obtained by the network with discriminative embeddings to build a cross-channel semantic association matrix enriched with additional information. This approach effectively supervises the learning of hash codes and improves the arrangement of semantic information across various modalities. Extensive empirical findings Indicate that on two commonly utilized datasets, our proposed CCUH method surpasses existing leading methods in cross-modal search tasks.