The human voice contains essential paralinguistic information that is used in many voice recognition applications. From a given voice dataset, gender prediction is considered a pivotal part to be classified using the speech data. This analysis is useful in gender recognition, speaker identification, and the detection of emotions. Several studies have been working on creating an automated system for gender categorization using speech datasets without prior knowledge of the speaker’s gender. The main focus of the study is extracting gender-specific traits from the speech data and creating a model to capture and train on these features. Machine Learning (ML) and Deep Learning (DL), especially Convolution Neural Networks (CNNs), have shown significant performance in this field. The study focused on increasing the accuracy of the models. This study investigates the different CNN models implemented on gender prediction using speech data containing voice spectrograms as input and compares their effectiveness to find the best DL technique. A publicly available dataset is used containing vocal features of both genders. It contains 3,168 voice recordings of male and female speakers, containing 20 acoustic features. Three CNN models were tested on this dataset, and the three-layer CNN model fared the best, with an accuracy of 98.21% and a loss of 5.64%. This research improves the performance of deep learning models in voice detection and conducts a comparative analysis among various models, ultimately determining the most effective model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Deep Learning Architecture for Accurate Gender Classification Using Vocal Data

  • Khushi Anand,
  • Bhawna Jain,
  • Ananya Verma,
  • Anushka Gupta,
  • Niharika Chhabra

摘要

The human voice contains essential paralinguistic information that is used in many voice recognition applications. From a given voice dataset, gender prediction is considered a pivotal part to be classified using the speech data. This analysis is useful in gender recognition, speaker identification, and the detection of emotions. Several studies have been working on creating an automated system for gender categorization using speech datasets without prior knowledge of the speaker’s gender. The main focus of the study is extracting gender-specific traits from the speech data and creating a model to capture and train on these features. Machine Learning (ML) and Deep Learning (DL), especially Convolution Neural Networks (CNNs), have shown significant performance in this field. The study focused on increasing the accuracy of the models. This study investigates the different CNN models implemented on gender prediction using speech data containing voice spectrograms as input and compares their effectiveness to find the best DL technique. A publicly available dataset is used containing vocal features of both genders. It contains 3,168 voice recordings of male and female speakers, containing 20 acoustic features. Three CNN models were tested on this dataset, and the three-layer CNN model fared the best, with an accuracy of 98.21% and a loss of 5.64%. This research improves the performance of deep learning models in voice detection and conducts a comparative analysis among various models, ultimately determining the most effective model.