ExARN: Target Speaker Extraction with Attentive Recurrent Networks
摘要
Target Speaker Extraction (TSE) isolates target speech from multi-speaker recordings, aiding applications like speaker verification. A significant challenge lies in effectively leveraging enrollment utterances to favor the target speaker, which, if not addressed properly, can result in speaker misidentification (referred to as speaker confusion, SC), leading to adverse user experiences. Although previous studies have investigated methods like concatenation and attention mechanisms to merge speaker embeddings from enrollment with mixed speech, these approaches have not fully capitalized on the potential of enrollment utterances. To overcome these limitations, We introduce ExARN, a novel TSE architecture with a custom self-attention mechanism and RNNs for better speaker embedding integration and extended sequence modeling. ExARN demonstrates unparalleled performance on benchmark datasets, significantly improving TSE.