<p>This paper introduces a novel speaker embedding loss function designed to improve the performance of end-to-end neural diarization (EEND) systems by enhancing speaker discrimination. Unlike previous methods that require additional speaker embedding networks or pre-training on large-scale speaker-labeled datasets, the proposed approach derives speaker-wise embeddings directly from frame-level encoder outputs and ground truth speaker labels. By avoiding external embedding networks, the architecture remains simple and efficient while still facilitating effective learning of speaker characteristics within a unified framework. The proposed loss incorporates cosine similarity to maximize the inter-speaker embedding distance, encouraging the model to learn discriminative embeddings for each speaker, and is integrated with a permutation-free loss to resolve label ambiguity during training. Experimental evaluations were conducted on both simulated and real-world datasets, including LibriSpeech and CALLHOME. The results demonstrate that the proposed speaker embedding loss significantly improves diarization accuracy, achieving a 20.8% relative reduction in Diarization Error Rate (DER) on the two-speaker LibriSpeech dataset (4.99% <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\rightarrow\)</EquationSource> <EquationSource Format="MATHML"><math> <mo stretchy="false">→</mo> </math></EquationSource> </InlineEquation> 3.95%). Consistent performance gains were also observed in five-speaker and CALLHOME evaluation settings. These findings underscore the potential of the proposed speaker embedding loss as a lightweight yet effective addition to EEND systems for improved diarization in overlapping and conversational speech scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker embedding loss for end-to-end speaker diarization without external embedding networks

  • Jaehee Jung,
  • Wooil Kim

摘要

This paper introduces a novel speaker embedding loss function designed to improve the performance of end-to-end neural diarization (EEND) systems by enhancing speaker discrimination. Unlike previous methods that require additional speaker embedding networks or pre-training on large-scale speaker-labeled datasets, the proposed approach derives speaker-wise embeddings directly from frame-level encoder outputs and ground truth speaker labels. By avoiding external embedding networks, the architecture remains simple and efficient while still facilitating effective learning of speaker characteristics within a unified framework. The proposed loss incorporates cosine similarity to maximize the inter-speaker embedding distance, encouraging the model to learn discriminative embeddings for each speaker, and is integrated with a permutation-free loss to resolve label ambiguity during training. Experimental evaluations were conducted on both simulated and real-world datasets, including LibriSpeech and CALLHOME. The results demonstrate that the proposed speaker embedding loss significantly improves diarization accuracy, achieving a 20.8% relative reduction in Diarization Error Rate (DER) on the two-speaker LibriSpeech dataset (4.99% \(\rightarrow\) 3.95%). Consistent performance gains were also observed in five-speaker and CALLHOME evaluation settings. These findings underscore the potential of the proposed speaker embedding loss as a lightweight yet effective addition to EEND systems for improved diarization in overlapping and conversational speech scenarios.