Feature Embedding Representation for Unsupervised Speaker Diarization in Telephone Calls
摘要
Speaker diarization aims to segment an audio recording, where different speakers are involved, into speech segments based on the speaker's identity. This study proposes a feature embedding representation for unsupervised speaker diarization in case of telephone conversations. First, the speech waveform is segmented into speech and non-speech segments using a voice activity detection algorithm. Then, a deep autoencoder performs feature embedding representation of the speech segments which are fed to the K-means algorithm to perform the clustering. Indeed, in the case of telephone calls speech data is often unequally shared between the two speakers, and an accurate feature representation helps to overcome limitations related to the varying width across clusters of the K-means algorithm. Experiments carried out on the CALLHOME dataset show the positive impact of the embedding representation in terms of Diarization Error Rate (DER) with an improvement of about 56%.