In recent years, significant progress in large language models has been driven by advancements in training strategies, instruction tuning, and the scaling of datasets. The strength of contemporary generative models lies in their ability to be guided through prompts. Contextual information, a type of prompt that generally exists in daily conversation, is essential for enhancing speech recognition in environments where speech is ambiguous or unclear. Inspired by this, we propose and evaluate two methods for our proposed Contextual Biasing Prompt-driven Speech Recognition (CB-PSR), which incorporates contextual biasing prompts as auxiliary information: 1) Early Prompting via Concatenation in Different Dimensions, and 2) Mid Prompting via the Cross-Attention Mechanism. Extensive experiments on the LibriSpeech dataset demonstrate the effectiveness of our proposed framework in enhancing speech recognition accuracy in spoken language understanding, with the Mid Prompting method performing the best. Additionally, the framework’s versatility is highlighted by its superior performance across various prompt biasing scenarios. This study establishes a benchmark and provides valuable insights for prompt-driven speech recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CB-PSR: Adaptive Contextual Biasing for Prompt-Driven Speech Recognition

  • Hongli Yang,
  • Ziyuan Chen,
  • Haowen Yin,
  • Hao Huang

摘要

In recent years, significant progress in large language models has been driven by advancements in training strategies, instruction tuning, and the scaling of datasets. The strength of contemporary generative models lies in their ability to be guided through prompts. Contextual information, a type of prompt that generally exists in daily conversation, is essential for enhancing speech recognition in environments where speech is ambiguous or unclear. Inspired by this, we propose and evaluate two methods for our proposed Contextual Biasing Prompt-driven Speech Recognition (CB-PSR), which incorporates contextual biasing prompts as auxiliary information: 1) Early Prompting via Concatenation in Different Dimensions, and 2) Mid Prompting via the Cross-Attention Mechanism. Extensive experiments on the LibriSpeech dataset demonstrate the effectiveness of our proposed framework in enhancing speech recognition accuracy in spoken language understanding, with the Mid Prompting method performing the best. Additionally, the framework’s versatility is highlighted by its superior performance across various prompt biasing scenarios. This study establishes a benchmark and provides valuable insights for prompt-driven speech recognition.