A Coarse and Fine Grained Masking Approach for Video-Grounded Dialogue
摘要
The task of Video-Grounded Dialogue involves developing a multimodal chatbot capable of answering sequential questions from humans regarding video content, audio, captions and dialog history. Although existing approaches utilizing deep learning models have demonstrated impressive performance, their success often relies on fusion of multimodal information in the limited datasets rather than understanding the interactions and dependencies among individual modalities such as video-audio, caption, dialog history, question or answer. In this paper, we present CFM (Coarse and Fine Grained Masking), a novel approach based on the pre-training model GPT2, aiming at enhancing cross-modal understanding among individual modalities in video-grounded dialogue. CFM achieves this by employing distinct coarse-grained and fine-grained masking strategies to differentiate various inputs, including video-audio, caption, dialog history, question and answer. Furthermore, we improve GPT2 model to strengthen its ability to integrate video-audio with text information effectively by incorporating multimodal feedforward network. Through extensive experiments on the Audio Visual Scene-Aware Dialog (AVSD) datasets, our proposed approach demonstrates promising performance, highlighting the benefits of our method in effectively figuring out dependencies and interactions among individual modalities.