Advances in visual, acoustic, and language processing technologies have led to a growing focus on multimodal information processing. Event extraction, as a branch of natural language processing, focuses mainly on the textual modality. Most existing multimodal event extraction methods have introduced visual information to text, while ignoring acoustic data. In this paper, we propose a novel model, TEE-CS (Text-oriented Event Extraction Combining Spectrogram), which combines acoustic modalities to enhance text event extraction. TEE-CS involves the construction of a multimodal dataset comprising text-spectrogram pairs. The speech of the sentence is synthesized via generative modeling, followed by extracting the Mel spectrogram of the speech as an audio feature. Subsequently, a two-stream multimodal model is designed to fuse the acoustic clues in the text and learn through an incremental training strategy. Finally, events are extracted from the multimodal representation. Experimental results on the ACE2005 English corpus demonstrate that combining speech clues enhances the performance of event extraction in comparison with the baseline models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-Oriented Event Extraction Combining Spectrogram

  • Liyao Xing,
  • Zhongqing Wang,
  • Peifeng Li

摘要

Advances in visual, acoustic, and language processing technologies have led to a growing focus on multimodal information processing. Event extraction, as a branch of natural language processing, focuses mainly on the textual modality. Most existing multimodal event extraction methods have introduced visual information to text, while ignoring acoustic data. In this paper, we propose a novel model, TEE-CS (Text-oriented Event Extraction Combining Spectrogram), which combines acoustic modalities to enhance text event extraction. TEE-CS involves the construction of a multimodal dataset comprising text-spectrogram pairs. The speech of the sentence is synthesized via generative modeling, followed by extracting the Mel spectrogram of the speech as an audio feature. Subsequently, a two-stream multimodal model is designed to fuse the acoustic clues in the text and learn through an incremental training strategy. Finally, events are extracted from the multimodal representation. Experimental results on the ACE2005 English corpus demonstrate that combining speech clues enhances the performance of event extraction in comparison with the baseline models.