错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker diarization based on X vector extracted from time-delay neural networks (TDNN) using agglomerative hierarchical clustering in noisy environment

  • K. V. Aljinu Khadar,
  • R. K. Sunil Kumar,
  • V. V. Sameer

摘要

This paper introduces a speaker diarization system using speaker embedding parameters, specifically the x-vector. By incorporating auto-correlated MFCC features for x-vector extraction using a pre-trained time delay neural network, the system exhibits enhanced adaptability to noise variations. Speaker clustering is accomplished through agglomerative clustering with PLDA scoring as the distance metric, making the system particularly valuable for potential speaker identification, especially in forensic applications. The system’s noise adaptability is thoroughly evaluated by integrating various types of noise such as Red, Pink, and white noise across a wide range of signal-to-noise ratios (20 dB to − 20 dB). Additionally, the system’s performance is comprehensively assessed by varying speech duration and adjusting the number of speakers, highlighting its robustness and effectiveness in real-world scenarios.