Speaker recognition has proved to be an important field of research in the context of automatic speech recognition systems contributing applications, ranging from security authentication to speech analysis. In this paper, optimizations of speaker recognition systems by advanced feature extraction techniques are discussed in particular with the Gujarati dialects. Gujarati is one of the dialects majorly spoken in western India. These dialects, which are difficult and distinguishable due to variations in phonetics, intonation, and articulation make this attempt distinctively challenging when implementing speaker recognition techniques. In this study, we would assess various feature extraction methods; this includes Mel-Frequency Cepstral Coefficients (MFCC) to Linear Predictive Coding (LPC) and Fast Fourier Transform (FFT) is compared in detail with Discrete Wavelet Transform (DWT), Perceptual Linear Prediction (PLP), and Relative Spectral (RASTA-PLP) methods and their impacts on speakers of those dialects. Accuracy results indicate a significant improvement when dialect-specific features are incorporated into the system, thus opening a new field toward more robust and adaptive speaker recognition systems. The results of this work will therefore be useful for more accurate and reliable construction of speaker recognition systems for Gujarati as well as other underrepresented languages. In bridging this gap in dialect-specific research in speaker recognition, it contributes to increasing the availability and accuracy of automatic speaker recognition systems across varying linguistic conditions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Speaker Recognition Through Feature Extraction Techniques: A Focus on Gujarati Dialects

  • Meera M. Shah,
  • Hiren Kavathiya

摘要

Speaker recognition has proved to be an important field of research in the context of automatic speech recognition systems contributing applications, ranging from security authentication to speech analysis. In this paper, optimizations of speaker recognition systems by advanced feature extraction techniques are discussed in particular with the Gujarati dialects. Gujarati is one of the dialects majorly spoken in western India. These dialects, which are difficult and distinguishable due to variations in phonetics, intonation, and articulation make this attempt distinctively challenging when implementing speaker recognition techniques. In this study, we would assess various feature extraction methods; this includes Mel-Frequency Cepstral Coefficients (MFCC) to Linear Predictive Coding (LPC) and Fast Fourier Transform (FFT) is compared in detail with Discrete Wavelet Transform (DWT), Perceptual Linear Prediction (PLP), and Relative Spectral (RASTA-PLP) methods and their impacts on speakers of those dialects. Accuracy results indicate a significant improvement when dialect-specific features are incorporated into the system, thus opening a new field toward more robust and adaptive speaker recognition systems. The results of this work will therefore be useful for more accurate and reliable construction of speaker recognition systems for Gujarati as well as other underrepresented languages. In bridging this gap in dialect-specific research in speaker recognition, it contributes to increasing the availability and accuracy of automatic speaker recognition systems across varying linguistic conditions.