Bivariate Empirical Mode Decomposition of Speech Signals for Disordered Voices Assessment
摘要
The acoustic analysis of the speech signal offers a privileged means for clinical assessment of the quality of the voice with a view to diagnosis and quantitative documentation of pathologies of the vocal folds. The major challenge of the acoustic analysis of the speech signal is to find reliable and accurate acoustic cues in order to determine the characteristics of the voice which provide information on the state of the speaker’s larynx. We propose bivariate empirical mode decomposition (BEMD) for multivariate analysis of vocal dysperiodicities for disordered voices assessment. The BEMD is an extension of the conventional empirical mode decomposition devoted to the decomposition of bivariate signals. The BEMD is applied jointly in the time and frequency domains. The acoustic cues computed from the complex intrinsic mode functions (IMFs) in the time and frequency domains are used as predictor variables of the scores of the perceived hoarseness. The proposed method is tested on datasets comprising Spanish sustained vowels /a/, English sustained vowels /a/ and English sentences produced by healthy and pathological speakers. The results indicate that the proposed approach outperforms reference methods which are multi-band generalized variogram-based vocal dysperiodicies analysis method, higher-order statistics-based method and correlation-based method in terms of correlation of the acoustic cue with the scores of perceived hoarseness. The achieved improvements exceed 29% for Spanish sustained vowels, 2.6% for English sustained vowels and 21% for English sentences.