MemXNet: A Multimodal Approach for Video Memorability Prediction
摘要
Video memorability prediction is a challenging task with applications in numerous domains, such as advertising, social media, and education. This paper proposes a multimodal deep learning approach to video memorability prediction that leverages both visual and textual information. A novel language-guided frame extraction process is also proposed to extract representative frames from the video. The proposed architecture, MemXNet, leverages spatiotemporal features from input video, semantic embeddings from the text description of the video, and visual features from the representative frames of the video to obtain a holistic representation for video memorability prediction. The performance of the model is evaluated on the VideoMem dataset and the results are presented. In-depth ablation studies are conducted to examine the contribution of each modality to the overall memorability score. This study contributes to the growing body of research on multimodal deep learning for video analysis and paves the way for future research in this area.