Prompt-Based and Modality-Semantic Enhanced Multimodal Recommendation
摘要
In contrast to traditional models relying exclusively on user-item interactions, multimodal recommendations also utilize multimodal item data to improve recommendation performance. Multimodal recommender systems focus on learning representations for both users and items. In practice, items tend to come with multimodal information, including text, images, and audio. These modalities serve as rich auxiliary information that can help item embeddings capture more semantic information. However, user embeddings are general-purpose representations that learn interests from modality-specific items equally through GCN, which results in limited expressiveness of user representations. Motivated by this, a new method is proposed, named prompt-based and modality-semantic enhanced multimodal recommendation, which introduces prompt embeddings and leverages user and item embeddings to filter out noisy prompt information. Moreover, two fusion strategies are designed to help user embeddings effectively incorporate prompt information. For multimodal information fusion, contrastive learning is employed to better integrate the multimodal semantic information extracted from the item-item structural graph into item representations, by aligning the semantic information before and after fusion to enhance the quality of representation learning. Extensive experiments on three Amazon datasets verified the effectiveness of the proposed components.