Exploring CNNs for Text Independent Speaker Verification for Diverse Indian Scenarios
摘要
The process of automatically identifying a speaker from a particular voice utterance is known as speaker recognition (SR). The SR approaches used in India are primarily taken from outside sources. Although the literature in publication demonstrates excellent performance, with EERs generally less than 2%, these findings are obtained in carefully controlled foreign contexts, like one speaker functioning in a quiet office setting. Trying to use these technologies in India’s vast and varied environments presents challenges. The fact that these technologies have not been developed with Indian operational conditions in mind is particularly noteworthy. Variations in dialects, accents, multispeaker scenarios, multilingualism, numerous channels, and sensors still need to be addressed. Deep learning is increasingly being used in speaker recognition as a competitive alternative to i-vectors. Convolutional neural networks (CNNs) have recently demonstrated encouraging performance when given raw voice samples directly. This work presents a CNN architecture for an Indian setting and compares the performance of CNN with Sincnet and VGG 16 architecture on the IIT-GMV data corpus.