Unlike traditional recommendation systems that rely solely on historical user-item interactions for recommendations, multimodal recommendation systems integrate multimodal features of items to enhance the accuracy of recommended results. In multimodal recommendation systems, generating multimodal embeddings is crucial for accurately recommending products of interest to target users. In many early models, the generation of multimodal embeddings lacked rigor and contained significant modal noise, resulting in underutilization of the effective information from the modalities. Therefore, we propose the Multi-modal Information Multi-Angle Mining (MIMM) for multimedia recommendation method to address this issue. Specifically, we have designed a novel multimodal embedding generator aimed at better aligning with historical interaction data and multimodal information. It consists of two main components: a similarity matrix that captures relationships between items, and a semantic fusion unit that integrates multimodal information. These components effectively capture the semantic essence of multimodal data, filter out noise, and enhance the precision of recommendations. In addition, we have developed a multimodal latent semantic miner that can identify unique semantics from latent modalities that exist outside of common modalities, emphasizing the independence of each modality. We validated the effectiveness of the model on three real-world public datasets, and found that it is better in recommendation performance than all existing advanced methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Information Multi-angle Mining for Multimedia Recommendation

  • Yijie Zhu,
  • Mingyong Li

摘要

Unlike traditional recommendation systems that rely solely on historical user-item interactions for recommendations, multimodal recommendation systems integrate multimodal features of items to enhance the accuracy of recommended results. In multimodal recommendation systems, generating multimodal embeddings is crucial for accurately recommending products of interest to target users. In many early models, the generation of multimodal embeddings lacked rigor and contained significant modal noise, resulting in underutilization of the effective information from the modalities. Therefore, we propose the Multi-modal Information Multi-Angle Mining (MIMM) for multimedia recommendation method to address this issue. Specifically, we have designed a novel multimodal embedding generator aimed at better aligning with historical interaction data and multimodal information. It consists of two main components: a similarity matrix that captures relationships between items, and a semantic fusion unit that integrates multimodal information. These components effectively capture the semantic essence of multimodal data, filter out noise, and enhance the precision of recommendations. In addition, we have developed a multimodal latent semantic miner that can identify unique semantics from latent modalities that exist outside of common modalities, emphasizing the independence of each modality. We validated the effectiveness of the model on three real-world public datasets, and found that it is better in recommendation performance than all existing advanced methods.