Auxiliary Context Module and Weighted Multihead Fusion for Multimodal Intent Recognition
摘要
Intent recognition is a crucial task in natural language understanding. In the real world, discerning human intentions requires integrating information from speech, facial expressions, and gestures, among other modalities. Multimodal intent recognition can combine data from diverse modalities to interpret human language and behavior. However, most existing methods struggle to process large volumes of redundant, unaligned multimodal sequential data, leading to inefficiencies in modeling multimodal fusion within such unaligned datasets. The proposed method introduces an Auxiliary Context Module (ACM). In particular, the ACM uses each modality's utterance-level representations as a global multimodal context, which interacts with local unimodal information to improve each other. This method offers better performance than earlier local-local cross-modal interaction strategies while also avoiding the quadratic scaling penalty. The Weighted Multihead Fusion Network (WMF) further refines fusion results through gated neural networks and multi-head attention systems. In experiments, the proposed method compares to the most advanced methods, significant performance improvements have been achieved on two datasets. Additionally, ablation experiments confirm the significant contributions of the ACM module and the WMF method in enhancing modal feature representation and improving intent recognition performance.