Explicit-implicit feature modeling in drama videos: dataset and benchmark
摘要
Understanding complicated scenes in movies or dramas is a foundational yet challenging task in video analysis, with significant implications for cultural understanding and intelligent video applications. While recent multimodal learning approaches focus on explicit features (e.g., visual, audio, and text), they often neglect the implicit background knowledge, such as cultural, social, and historical contexts, which are pivotal for accurate video understanding. In this paper, we propose a new task, named Explicit-Implicit feature representation in Video Understanding (EIVU). To facilitate this task, we introduce the Explicit-Implicit Drama (EIDrama), a multimodal dataset based on traditional Chinese drama videos, comprising over one million frames sampled from hundreds of video hours. Besides the explicit visual and audio features, EIDrama also consists of abundant background information, enabling a comprehensive exploration of the interplay between explicit and implicit modalities. Additionally, we present EINet, a baseline model that aggregates explicit and implicit features using a local-global strategy. Extensive experiments demonstrate the effectiveness of EINet. We expect the proposed EIDrama and EINet as stepping stones towards bridging the gap between explicit-implicit modeling and fostering cross-cultural understanding in multimodal video analysis.