The process of automatically identifying a speaker from a particular voice utterance is known as speaker recognition (SR). The SR approaches used in India are primarily taken from outside sources. Although the literature in publication demonstrates excellent performance, with EERs generally less than 2%, these findings are obtained in carefully controlled foreign contexts, like one speaker functioning in a quiet office setting. Trying to use these technologies in India’s vast and varied environments presents challenges. The fact that these technologies have not been developed with Indian operational conditions in mind is particularly noteworthy. Variations in dialects, accents, multispeaker scenarios, multilingualism, numerous channels, and sensors still need to be addressed. Deep learning is increasingly being used in speaker recognition as a competitive alternative to i-vectors. Convolutional neural networks (CNNs) have recently demonstrated encouraging performance when given raw voice samples directly. This work presents a CNN architecture for an Indian setting and compares the performance of CNN with Sincnet and VGG 16 architecture on the IIT-GMV data corpus.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring CNNs for Text Independent Speaker Verification for Diverse Indian Scenarios

  • Satish Chikkamath,
  • S. R. Nirmala

摘要

The process of automatically identifying a speaker from a particular voice utterance is known as speaker recognition (SR). The SR approaches used in India are primarily taken from outside sources. Although the literature in publication demonstrates excellent performance, with EERs generally less than 2%, these findings are obtained in carefully controlled foreign contexts, like one speaker functioning in a quiet office setting. Trying to use these technologies in India’s vast and varied environments presents challenges. The fact that these technologies have not been developed with Indian operational conditions in mind is particularly noteworthy. Variations in dialects, accents, multispeaker scenarios, multilingualism, numerous channels, and sensors still need to be addressed. Deep learning is increasingly being used in speaker recognition as a competitive alternative to i-vectors. Convolutional neural networks (CNNs) have recently demonstrated encouraging performance when given raw voice samples directly. This work presents a CNN architecture for an Indian setting and compares the performance of CNN with Sincnet and VGG 16 architecture on the IIT-GMV data corpus.