Utilizing Convolutional Neural Networks and Mel Spectrograms for Indian Spoken Language Detection
摘要
The idea behind automatic language recognition is to identify the spoken language by analyzing specific attributes unique to each language. Various methods, ranging from straightforward K-NN models to sophisticated deep neural networks, have been developed for this purpose. The accuracy of these models has significantly improved due to advancements in classification algorithms and the computational capabilities of machines. This paper performs a comparative analysis of different approaches to automatic language recognition, exploring various features and classification algorithms to determine the most effective method for automatically identifying spoken languages. The study uses approximately 10 h of training data and 2 h of test audio data, obtained from All India Radio, covering seven languages: Bengali, Gujarati, Hindi, Marathi, Punjabi, Telugu, and Tamil. These languages were selected to represent the diverse linguistic regions of the country.