Unsupervised phoneme segmentation of continuous Arabic speech
摘要
The development of a speech recognition system for the Arabic language presents a significant challenge, mainly due to the limited availability of digital resources specific to this language. To achieve vocabulary-independent speech recognition, it is essential to split a given speech into smaller units known as phonemes or syllables. This process is what we call speech segmentation; it plays a crucial role in accurately recognizing and understanding speech patterns. Many speech segmentation techniques have been developed, relying on linguistic information such as phonetic transcription. However, for real-time systems, phonetic transcription is not always available, especially for low-resource languages like Arabic. In this paper, we addressed the problem of unsupervised segmentation for continuous Arabic speech based on two distinct approaches: spectral contrast and the first derivative of Mel Frequency Cepstrum Coefficients (