Enhancing Automatic Speaker Diarization for Marathi and Hindi Languages: Feature Extraction Techniques
摘要
Automatic speaker diarization is essential in numerous speech analysis programs, such as event recording, recognizing speakers, or sound search. In this study, researchers investigate the use of extraction of features approaches for automated speech diarization in Marathi and Hindi languages. To guarantee the resilience and correctness of the diarization approach, we diligently collected a large dataset of different examples of speech in each language. MFCCs to extract features, with average and standard deviation values of 2.4 and 6.68, respectively. MFCCs, which are well known for their efficacy in speech processing, are especially good at catching the subtleties of the voice of a person. MFCCs are analyze all frequency ranges of a spoken stream and offer an accurate depiction of phonological material, which is critical for differentiating different speakers. Our technique intends to improve speaker diarization accuracy by utilizing complex extraction of features methods, ultimately contributing to the improvement of speech-processing technology for neglected dialects like Marathi and Hindi.