<p>In natural language understanding, intent recognition stands out as a crucial task that has drawn significant attention. While previous research focuses on intent recognition using task-specific unimodal data, real-world scenarios often involve human intents expressed through various ways, including speech, tone of voice, facial expressions, and actions. This prompts research into integrating multimodal information to more accurately identify human intent. However, existing intent recognition studies often fuse textual and non-textual modalities without considering their quality gap. The gap in feature quality across different modalities hinders the improvement of the model’s performance. To address this challenge, we propose a multimodal intent recognition model to enhance non-textual modality features. Specifically, we enrich the semantics of non-textual modalities by replacing redundant information through text-guided cross-modal attention. Additionally, we introduce a text-centric adaptive fusion gating mechanism to capitalize on the primary role of text modality in intent recognition. Extensive experiments on two multimodal task datasets show that our proposed model performs better in all metrics than state-of-the-art multimodal models. The results demonstrate that our model efficiently enhances non-textual modality features and fuses multimodal information, showing promising potential for intent recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal intent recognition based on text-guided cross-modal attention

  • Zhengyi Li,
  • Junjie Peng,
  • Xuanchao Lin,
  • Zesu Cai

摘要

In natural language understanding, intent recognition stands out as a crucial task that has drawn significant attention. While previous research focuses on intent recognition using task-specific unimodal data, real-world scenarios often involve human intents expressed through various ways, including speech, tone of voice, facial expressions, and actions. This prompts research into integrating multimodal information to more accurately identify human intent. However, existing intent recognition studies often fuse textual and non-textual modalities without considering their quality gap. The gap in feature quality across different modalities hinders the improvement of the model’s performance. To address this challenge, we propose a multimodal intent recognition model to enhance non-textual modality features. Specifically, we enrich the semantics of non-textual modalities by replacing redundant information through text-guided cross-modal attention. Additionally, we introduce a text-centric adaptive fusion gating mechanism to capitalize on the primary role of text modality in intent recognition. Extensive experiments on two multimodal task datasets show that our proposed model performs better in all metrics than state-of-the-art multimodal models. The results demonstrate that our model efficiently enhances non-textual modality features and fuses multimodal information, showing promising potential for intent recognition.