错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Gender and Age Extraction from Audio Signal Using Convolutional Neural Network, MFCC and Spectrogram

  • Fazia Karaoui,
  • Rachida Djeradi,
  • Amar Djeradi

摘要

As technology continues to advance and artificial intelligence grows, accurately and efficiently determining the age and gender of a speaker from speech signals paves the way for developing applications across diverse sectors. Our work aims to automatically determine the age and gender of speakers from audio data. In this study, we used two distinct databases, presenting the characteristics and content of each: the Mozilla Common Voice dataset, we used the Arabic language subset and we designed a database for Algerian Arabic dialects. Our approach to classifying age and gender from voice involves combining two methodologies. Firstly, we utilize a set of 20 MFCC coefficients extracted from each audio file. Secondly, we employ the log mel-spectrogram to treat audio classification as an image classification task, leveraging CNN models. We have selected a serial model to classify the gender and age of speakers. This architecture provides the opportunity to examine gender and age classification problems independently. Optimal results are obtained; our model achieved an accuracy of 99.44% for the gender recognition regarding Mozilla dataset and 97.83 regarding Algerian dataset, which is the most notable achievements in gender classification task for Arabic speakers has been achieved, as documented in the literature. For the age recognition, the accuracy reached 83.40% for male speakers and 79.6% for female speakers using the Algerian dataset.