Audio Visual Scene-Aware Dialog (AVSD) is to accurately generate the answer of current question based on multimodal information, such as video, audio, and dialog history. Multimodal data collected in the real world is often incomplete, often occurring partial modal data loss, which poses a significant challenge to AVSD. To address this issue, we propose Multimodal Prompt Learning (MPL) for AVSD, a novel approach that preserves the input format of missing data by setting up virtual modalities. We design various types of learnable prompts and MPL can optimize prompt parameters based on the different modal data loss situations, then add them before the input sequence. Through multi-level prompt delivery, MPL exploits the inherited prompt information from the previous layer to learn more effective instructions for each level of prompts, which enhances the context understanding and information compensation. As a result, only less than 1% additional parameters are needed to adapt to the pre-trained model, but it relieves the training problem caused by missing partial modalities in video-grounded dialogue. We conduct extensive experiments on the cases of missing some video-audio or dialog history, and discuss the robustness of our proposed method compared to the baselines. In order to further explore our MPL under different resource settings, we also investigate the number of samples and the number of layers for prompts. A large amount of experimental results demonstrate that our MPL achieves competitive results on the DSTC7@ AVSD and DSTC8@ AVSD datasets even when the training set is modality-deficient.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Prompt Learning for Audio Visual Scene-Aware Dialog

  • Feifei Xu,
  • Fumiaoyue Jia,
  • Wang Zhou

摘要

Audio Visual Scene-Aware Dialog (AVSD) is to accurately generate the answer of current question based on multimodal information, such as video, audio, and dialog history. Multimodal data collected in the real world is often incomplete, often occurring partial modal data loss, which poses a significant challenge to AVSD. To address this issue, we propose Multimodal Prompt Learning (MPL) for AVSD, a novel approach that preserves the input format of missing data by setting up virtual modalities. We design various types of learnable prompts and MPL can optimize prompt parameters based on the different modal data loss situations, then add them before the input sequence. Through multi-level prompt delivery, MPL exploits the inherited prompt information from the previous layer to learn more effective instructions for each level of prompts, which enhances the context understanding and information compensation. As a result, only less than 1% additional parameters are needed to adapt to the pre-trained model, but it relieves the training problem caused by missing partial modalities in video-grounded dialogue. We conduct extensive experiments on the cases of missing some video-audio or dialog history, and discuss the robustness of our proposed method compared to the baselines. In order to further explore our MPL under different resource settings, we also investigate the number of samples and the number of layers for prompts. A large amount of experimental results demonstrate that our MPL achieves competitive results on the DSTC7@ AVSD and DSTC8@ AVSD datasets even when the training set is modality-deficient.