Multimodal dialogue systems aim to process multimodal input information such as text and images simultaneously and then generate coherent and meaningful output dialogue responses. Although current methods have achieved notable progress, they have shortcomings in the following aspects: 1) Insufficient modeling of relationships between multimodal semantic elements, especially ignoring the interaction between the ordinal information in text and the position information of images. 2) Ineffectively integrating information from various modalities in multimodal conversations. To address these limitations, this paper proposes a framework for a multimodal dialogue system. We integrate both ordinal words from text and the position information of images as ordinal information to strengthen the interaction between multimodal semantic elements and to improve the understanding of user intent. In particular, in this framework, we incorporated the self-attention mechanism into the multimodal factorized bilinear pooling (MFB) method, which captures long-range dependencies between modalities, enhancing the effectiveness of multimodal data fusion. Finally, we conducted comprehensive experiments on public multimodal dialogue datasets (MMConv and MMD) and demonstrated that our proposed framework outperforms other methods in the experiments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ordinal and Position Enhance the Framework of the Multimodal Dialogue System

  • Longmei Xu,
  • Fei Chen,
  • Jinyu Li

摘要

Multimodal dialogue systems aim to process multimodal input information such as text and images simultaneously and then generate coherent and meaningful output dialogue responses. Although current methods have achieved notable progress, they have shortcomings in the following aspects: 1) Insufficient modeling of relationships between multimodal semantic elements, especially ignoring the interaction between the ordinal information in text and the position information of images. 2) Ineffectively integrating information from various modalities in multimodal conversations. To address these limitations, this paper proposes a framework for a multimodal dialogue system. We integrate both ordinal words from text and the position information of images as ordinal information to strengthen the interaction between multimodal semantic elements and to improve the understanding of user intent. In particular, in this framework, we incorporated the self-attention mechanism into the multimodal factorized bilinear pooling (MFB) method, which captures long-range dependencies between modalities, enhancing the effectiveness of multimodal data fusion. Finally, we conducted comprehensive experiments on public multimodal dialogue datasets (MMConv and MMD) and demonstrated that our proposed framework outperforms other methods in the experiments.