Optimizing Speaker Recognition Through Feature Extraction Techniques: A Focus on Gujarati Dialects
摘要
Speaker recognition has proved to be an important field of research in the context of automatic speech recognition systems contributing applications, ranging from security authentication to speech analysis. In this paper, optimizations of speaker recognition systems by advanced feature extraction techniques are discussed in particular with the Gujarati dialects. Gujarati is one of the dialects majorly spoken in western India. These dialects, which are difficult and distinguishable due to variations in phonetics, intonation, and articulation make this attempt distinctively challenging when implementing speaker recognition techniques. In this study, we would assess various feature extraction methods; this includes Mel-Frequency Cepstral Coefficients (MFCC) to Linear Predictive Coding (LPC) and Fast Fourier Transform (FFT) is compared in detail with Discrete Wavelet Transform (DWT), Perceptual Linear Prediction (PLP), and Relative Spectral (RASTA-PLP) methods and their impacts on speakers of those dialects. Accuracy results indicate a significant improvement when dialect-specific features are incorporated into the system, thus opening a new field toward more robust and adaptive speaker recognition systems. The results of this work will therefore be useful for more accurate and reliable construction of speaker recognition systems for Gujarati as well as other underrepresented languages. In bridging this gap in dialect-specific research in speaker recognition, it contributes to increasing the availability and accuracy of automatic speaker recognition systems across varying linguistic conditions.