Movie Retrieval Systems Using Genre-Guided Multimodal Learning Techniques
摘要
Given the rapid development of streaming platforms in recent years, movie retrieval has become an important and challenging research topic. However, compared to short video clips, the complex narrative structure of movies makes it difficult for models to understand audiovisual content, resulting in suboptimal movie representation learning. To address this issue, we propose a multimodal alignment-based movie retrieval method that integrates audiovisual features with textual information, such as plot summaries and genre descriptions, to more accurately reflect the themes, plots, and styles of movies. Our method trains audiovisual encoders to align with pre-trained text encoders, thereby extracting highly semantically rich movie audiovisual representations. Validated on the LVU benchmark [6], our approach demonstrates superior performance in movie retrieval tasks, effectively retrieving movie clips based on user descriptions or genre tags. Additionally, we have designed a visualization interface for the retrieval results, enhancing the user experience and showcasing the advantages of our method. Our demo video can be accessed through the following link: https://youtu.be/dIoFOm0Sy48