Speaker embedding loss for end-to-end speaker diarization without external embedding networks
摘要
This paper introduces a novel speaker embedding loss function designed to improve the performance of end-to-end neural diarization (EEND) systems by enhancing speaker discrimination. Unlike previous methods that require additional speaker embedding networks or pre-training on large-scale speaker-labeled datasets, the proposed approach derives speaker-wise embeddings directly from frame-level encoder outputs and ground truth speaker labels. By avoiding external embedding networks, the architecture remains simple and efficient while still facilitating effective learning of speaker characteristics within a unified framework. The proposed loss incorporates cosine similarity to maximize the inter-speaker embedding distance, encouraging the model to learn discriminative embeddings for each speaker, and is integrated with a permutation-free loss to resolve label ambiguity during training. Experimental evaluations were conducted on both simulated and real-world datasets, including LibriSpeech and CALLHOME. The results demonstrate that the proposed speaker embedding loss significantly improves diarization accuracy, achieving a 20.8% relative reduction in Diarization Error Rate (DER) on the two-speaker LibriSpeech dataset (4.99%