The rise of voice interface applications has renewed interest in improving the robustness of spoken language understanding(SLU). Many advances have come from end-to-end speech-language joint training, such as inferring semantics directly from speech signals and post-editing automatic speech recognition (ASR) output. Despite their performance achievements, these methods either suffer from the unavailability of a large number of paired error-prone ASR transcriptions and ground-truth annotations or are computationally costly. To mitigate these issues, we propose an ASR-robust pre-trained language model (ASRLM), which involves a generator generating simulated ASR transcriptions from ground-truth annotations and a sample-efficient discriminator distinguishing reasonable ASR errors from unrealistic ones. Experimental results demonstrate that ASRLM improves performance on a wide range of SLU tasks in the presence of ASR errors while saving 27% of the computation cost compared to baselines. Analysis also shows that our proposed generator is better than other simulation methods, including both BERT and GPT4-based, at simulating real-world ASR error situations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ASRLM: ASR-Robust Language Model Pre-training via Generative and Discriminative Learning

  • Qian Hu,
  • Xue Han,
  • Yiting Wang,
  • Yitong Wang,
  • Chao Deng,
  • Junlan Feng

摘要

The rise of voice interface applications has renewed interest in improving the robustness of spoken language understanding(SLU). Many advances have come from end-to-end speech-language joint training, such as inferring semantics directly from speech signals and post-editing automatic speech recognition (ASR) output. Despite their performance achievements, these methods either suffer from the unavailability of a large number of paired error-prone ASR transcriptions and ground-truth annotations or are computationally costly. To mitigate these issues, we propose an ASR-robust pre-trained language model (ASRLM), which involves a generator generating simulated ASR transcriptions from ground-truth annotations and a sample-efficient discriminator distinguishing reasonable ASR errors from unrealistic ones. Experimental results demonstrate that ASRLM improves performance on a wide range of SLU tasks in the presence of ASR errors while saving 27% of the computation cost compared to baselines. Analysis also shows that our proposed generator is better than other simulation methods, including both BERT and GPT4-based, at simulating real-world ASR error situations.