Temporal feature-based approaches for enhancing phoneme boundary detection and masking in speech
摘要
Automatic phoneme boundary detection is a key problem in speech processing and applications. The accurate phoneme segmentation in continuous speech contributes to the improvement of recognition quality. This study proposes an efficient temporal features-based approach for the automatic detection and masking of phoneme boundaries in speech signals. Changes in speech waveform during a well-spoken word can be used to identify phonemes boundaries. The signal level properties of the speech transitions in the waveform during the transformation from one phoneme to the next are analysed to establish the phoneme boundaries of a speech signal. The short time energy, pitch, zero-crossing rate, and Teager energy operator are used to determine the region of change from unvoiced to voiced and vice versa. The technique proposed in this paper enumerates empirical estimations observed with desired statistical properties for speech signal during phoneme transformation. The results of this study’s demonstrate that precise phoneme boundaries can be obtained by using the proposed method. The proposed method is applied on the TIMIT corpus and achieved the accuracy of 96.04% is accomplished within a tolerance scale of 10 ms.