For the cross-media retrieval, due to the “heterogeneous gap” problem between different modal data caused by the existence of the characterization of inconsistencies, it is difficult to measure the similarity directly. To solve this problem, we propose a Cross-media Correlation Computational Method for Multimodal Semantic Sparse Data (CCCM). CCCM supports five modal semantic sparse data, including text, image, video, audio, and 3D model. First, we propose a fine-grained cross-media correlation learning method that fuses cross-entropy and distribution differences. The data in the dataset is used for cross-media correlation learning through fine-grained segmentation, LSTM network, and loss function that combines cross-entropy and distribution difference. On this basis, most existing methods only consider the correlation analysis between different media data and ignore the sparse semantic problem existing in cross-media datasets. So we proposed a keyword-based KTF-IDF method to quantify the correlation between semantic tags. By conducting experiments on a large-scale cross-media dataset containing five modal data, compared with the existing common methods, the experimental results show that the correlation obtained by CCCM can be applied to multi-modal data cross-media. It can improve the accuracy by an average of 23% when it comes to media retrieval tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Media Correlation Computational Method for Multimodal Semantic Sparse Data

  • Jingdong Wang,
  • Kaidi Tian,
  • Xiaoyu Chang,
  • Fanqi Meng

摘要

For the cross-media retrieval, due to the “heterogeneous gap” problem between different modal data caused by the existence of the characterization of inconsistencies, it is difficult to measure the similarity directly. To solve this problem, we propose a Cross-media Correlation Computational Method for Multimodal Semantic Sparse Data (CCCM). CCCM supports five modal semantic sparse data, including text, image, video, audio, and 3D model. First, we propose a fine-grained cross-media correlation learning method that fuses cross-entropy and distribution differences. The data in the dataset is used for cross-media correlation learning through fine-grained segmentation, LSTM network, and loss function that combines cross-entropy and distribution difference. On this basis, most existing methods only consider the correlation analysis between different media data and ignore the sparse semantic problem existing in cross-media datasets. So we proposed a keyword-based KTF-IDF method to quantify the correlation between semantic tags. By conducting experiments on a large-scale cross-media dataset containing five modal data, compared with the existing common methods, the experimental results show that the correlation obtained by CCCM can be applied to multi-modal data cross-media. It can improve the accuracy by an average of 23% when it comes to media retrieval tasks.