Synth-ruSC: Construction and Validation of Synthetic Dataset to Solve the Problem of Keyword Spotting in Russian
摘要
This paper presents the experience of automatically creating a dataset of audio recordings of individual words in Russian. In the process of creating a dataset, technology of speech generation from given words and audio recordings with examples of pronunciation style is used taking into account the features of intonation and timbre of the voice. Due to the appearance after generation of low-quality audio recordings containing long pauses, distortion of words, as well as noise, a filtering function has been developed. It is based on identifying fragments with speech (voice activity detector), a set of rules for assessing the duration of selected fragments and the results of speech recognition models. As a result of using a set of generation and filtering tools, a synthetic dataset was created containing 61756 audio recordings with 39 words in Russian. The quality of the created data set was assessed using audio recordings from real people (896 audio from 23 people) based on a classification neural network model. It is shown that the implemented approach to creating a synthetic dataset can be used to solve the keyword spotting problem with a classification accuracy of 0.88 F1-macro for real audio recordings. The accuracy can be improved to 0.94 F1-macro when taking into account the differences between real and generated audio recordings.