Recent Methods and Algorithms in Speech Segmentation Tasks
摘要
The article addresses challenges in human-computer interaction through natural language, particularly in the context of collaborative conversations. The issue of overlapping audio data affects the accuracy of speech recognition and synthesis systems, especially in scenarios like meetings and negotiations. Emphasis is placed on the need for segregating and clustering speech from multiple speakers, highlighting challenges arising from diverse sound conditions in everyday life. The article then delves into the task of diarization, underscoring the importance of segmenting and processing speech data for effective voice control of devices. Subsequently, it explores the combination of GMM and i-vectors, as well as the evolution of approaches using deep learning, including convolutional and recurrent neural networks. Considering recent trends, the authors analyze the application of Transformers for handling long-term dependencies in data. The concluding section of the article provides a comprehensive overview and analysis of contemporary diarization methods, encompassing algorithms, error evaluation metrics, and descriptions of popular tools, with a focus on more modern approaches. This work constitutes a significant contribution to the field of speech diarization research, covering more current methods and trends compared to previous reviews in this domain.