<p>Speaker diarization is known for the “who spoke when?” problem which helps in understanding the audio content in a better and detailed manner. Although numerous efforts have been made towards efficient speaker diarization systems, they suffer from a common problem of overlapped speech. Overlapped speech makes extraction of the duration of an individual speaker a complex task in the presence of multiple speakers. The paper proposed a novel, Handled Overlap-Aware Refined Diarization (HOARD) framework to cater issue of overlapped speech in a multi-speaker diarization system. The proposed framework comprises Optimized Overlap-Aware Spectral Clustering (OOA-SC) and Overlapped Speakers’ Handling (OSH) module. The performance of the proposed framework is evaluated on benchmarked datasets; VoxConverse (dev and test), AMI, and DISPLACE2024 dataset. Results are obtained using the Diarization Error Rate (DER) metric as 8.76% for the dev set, 12.07% for the test set, 20.8%, and 29.0% for the AMI and DISPLACE 2024 dataset, respectively. Massive experiments are carried out to compare the effectiveness of the proposed model with baseline models. The proposed work has outperformed when compared with state-of-the-art approaches for both the VoxConverse and AMI datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Framework to Handle Overlapped Speech for Multiple Speakers in Speaker Diarization

  • Aishwarya Gupta,
  • Archana Purwar

摘要

Speaker diarization is known for the “who spoke when?” problem which helps in understanding the audio content in a better and detailed manner. Although numerous efforts have been made towards efficient speaker diarization systems, they suffer from a common problem of overlapped speech. Overlapped speech makes extraction of the duration of an individual speaker a complex task in the presence of multiple speakers. The paper proposed a novel, Handled Overlap-Aware Refined Diarization (HOARD) framework to cater issue of overlapped speech in a multi-speaker diarization system. The proposed framework comprises Optimized Overlap-Aware Spectral Clustering (OOA-SC) and Overlapped Speakers’ Handling (OSH) module. The performance of the proposed framework is evaluated on benchmarked datasets; VoxConverse (dev and test), AMI, and DISPLACE2024 dataset. Results are obtained using the Diarization Error Rate (DER) metric as 8.76% for the dev set, 12.07% for the test set, 20.8%, and 29.0% for the AMI and DISPLACE 2024 dataset, respectively. Massive experiments are carried out to compare the effectiveness of the proposed model with baseline models. The proposed work has outperformed when compared with state-of-the-art approaches for both the VoxConverse and AMI datasets.