Multimodal data fusion towards video summarization applications
摘要
The rapid rise of micro-content platforms has created a growing need for efficient video summarization techniques. This study presents a novel approach to video summarization by generating concise textual summaries through the fusion of textual information extracted from both audio and visual content. Unlike traditional methods that rely on direct video processing, the proposed approach leverages Natural Language Processing (NLP) and machine learning techniques to extract, integrate, and summarize textual data from video frames and audio transcripts. Specifically, a transformer-based model is employed, combining the global attention mechanism of Transformers with the sequential processing capabilities of LSTMs, enabling a more coherent and contextually rich summary. By integrating NLP techniques with transformer-based architectures, the effectiveness of video summarization is enhanced while optimizing computational efficiency. The proposed framework significantly reduces reliance on resource-intensive video training, making it a scalable and practical solution for diverse applications in video content analysis.