A Region Based Non-overlapping Reference Speech Estimation Method for Speaker Extraction
摘要
Speaker extraction is a technique that separates the target speech from multi-talker mixtures using a priori information about the target speaker, such as pre-enrolled reference speech. However, in real-world scenarios, the mixture speech is partially overlapped continuous long speech and obtaining a priori information is often challenging. Hence, we propose a framework to estimate the reference speech of participating speakers from non-overlapping input regions and extract target speech. To accurately estimate the regions, we adopt the idea of region proposal to generate multiple speech segment proposals of non-overlapping regions from the input speech mixtures. And then, we cluster these proposed segments into clusters to obtain the best reference speech of each speaker. We conduct experiment on simulated meeting-style test set with different overlap ratio based on LibriSpeech. The experimental results show that the region proposal method can achieve the best performance in speech extraction compared to other reference speech estimation methods.