Query-attentive video summarization: a comprehensive review
摘要
Since the last decade, the diverse applications of video summarization have gained increased attention, motivating researchers in the domain of computer vision to generate optimal and comprehensible video summaries. The main challenge in the research of video summarization is user perception and preference as humans are the ultimate consumers of generated summary. A single video summary cannot satisfy all users unless the summarization algorithm interacts with end users and adapts to their requirements. Conventional video summarization can not tackle the user requirements. This study explores various state-of-the-art techniques developed for generating user-intended video summaries, focusing on query-attentive video summarization. Query-attentive video summarization is a multi-modal summarization method that generates a video summary that satisfies the viewer’s requirements by taking input queries from the viewers. This paper discusses the fundamental aspects of query-attentive video summarization, tracing its progress and evolution over time. Contemporary approaches are explored in detail, highlighting developed techniques with advantages and limitations. Additionally, the article also studies publicly available datasets, including extensively utilized Query-Focused Video Summarization dataset, since these datasets ensure the validity and applicability of developed techniques. Evaluation metrics, which are essential tools for measuring performance and assessing user satisfaction are also studied and performance comparisons are presented. After investigating the domain of query-attentive video summarization, this article addresses the current research challenges and identifies potential future research objectives. This comprehensive review offers a complete guide for new researchers in the field of query-attentive video summarization, covering both existing and future real-time applications.