Speaker recognition is a method that assigns a distinct identity to each individual voice. For developing these systems generally, a large amount of data is required which should have a large number of speakers usually in thousands and a moderate amount of utterances per speaker. For under-resourced indo-aryan languages like Hindi such large datasets with a huge number of speakers are not freely available yet. This paper aims to study of how much data is required (in terms of number of speakers and amount of data in terms of hours) to train speaker recognition models for the under-resourced languages. To find the answer to this question, this paper presents the results of adapting the d-vector speaker verification model which is pre-trained on English data to the Hindi language. The results of the model are shown in terms of Equal Error Rate (EER) and Detection Cost Function (DCF) when the model is trained from scratch as well as after the adaption of the pre-trained English model. The EER of 4.46 \(\%\) was observed when the English model was finetuned using the Hindi data which was only around \(35\%\) of the English data in terms of a number of hours and the performance of this model is pretty close to the performance of the model trained in the English Language.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating Data Requirements for Hindi Speaker Recognition: A Comparative Study with English

  • Parth Khadse,
  • Sabyasachi Chandra,
  • Puja Bharati,
  • Debolina Pramanik,
  • G. Satya Prasad,
  • Aniket Aitawade,
  • Shyamal Kumar Das Mandal

摘要

Speaker recognition is a method that assigns a distinct identity to each individual voice. For developing these systems generally, a large amount of data is required which should have a large number of speakers usually in thousands and a moderate amount of utterances per speaker. For under-resourced indo-aryan languages like Hindi such large datasets with a huge number of speakers are not freely available yet. This paper aims to study of how much data is required (in terms of number of speakers and amount of data in terms of hours) to train speaker recognition models for the under-resourced languages. To find the answer to this question, this paper presents the results of adapting the d-vector speaker verification model which is pre-trained on English data to the Hindi language. The results of the model are shown in terms of Equal Error Rate (EER) and Detection Cost Function (DCF) when the model is trained from scratch as well as after the adaption of the pre-trained English model. The EER of 4.46 \(\%\) was observed when the English model was finetuned using the Hindi data which was only around \(35\%\) of the English data in terms of a number of hours and the performance of this model is pretty close to the performance of the model trained in the English Language.