Target Speaker Extraction (TSE) isolates target speech from multi-speaker recordings, aiding applications like speaker verification. A significant challenge lies in effectively leveraging enrollment utterances to favor the target speaker, which, if not addressed properly, can result in speaker misidentification (referred to as speaker confusion, SC), leading to adverse user experiences. Although previous studies have investigated methods like concatenation and attention mechanisms to merge speaker embeddings from enrollment with mixed speech, these approaches have not fully capitalized on the potential of enrollment utterances. To overcome these limitations, We introduce ExARN, a novel TSE architecture with a custom self-attention mechanism and RNNs for better speaker embedding integration and extended sequence modeling. ExARN demonstrates unparalleled performance on benchmark datasets, significantly improving TSE.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ExARN: Target Speaker Extraction with Attentive Recurrent Networks

  • Pengjie Shen,
  • Shulin He,
  • Xueliang Zhang

摘要

Target Speaker Extraction (TSE) isolates target speech from multi-speaker recordings, aiding applications like speaker verification. A significant challenge lies in effectively leveraging enrollment utterances to favor the target speaker, which, if not addressed properly, can result in speaker misidentification (referred to as speaker confusion, SC), leading to adverse user experiences. Although previous studies have investigated methods like concatenation and attention mechanisms to merge speaker embeddings from enrollment with mixed speech, these approaches have not fully capitalized on the potential of enrollment utterances. To overcome these limitations, We introduce ExARN, a novel TSE architecture with a custom self-attention mechanism and RNNs for better speaker embedding integration and extended sequence modeling. ExARN demonstrates unparalleled performance on benchmark datasets, significantly improving TSE.