Method of voice source coding with data compression based on the linear prediction model
摘要
The problem of voice source coding with data compression based on the linear prediction model is considered within the rapidly developing research area in the field of acoustic measurements—parameter analysis and estimation of an excitation signal inducing acoustic oscillations in a speaker’s vocal tract. With the application of the criterion of minimum average voice source power in speech production, the described problem is reduced to the real-time coding of a linear prediction error signal. A voice coding method that involves clipping of the linear prediction error was developed. The proposed method provides a means to avoid computationally intensive procedures for measuring the initial phase and fundamental frequency of a speech signal. An example of its technical implementation in soft real-time mode is considered. For a comparative effectiveness analysis of the proposed method and the widespread method of discrete cosine transform (DCT), a full-scale experiment was set up and conducted. It is shown that due to the reduction of data compression artifacts in the reconstructed speech signal, the accuracy of voice source coding via the developed method is one and a half to two times higher (as compared to the DCT method), and it is not necessary to detect vowel speech sounds and pauses in the speech signal. The obtained results can be used to develop new and upgrade existing systems and algorithms in the field of automatic speech processing and synthesis, mobile speech communication, artificial intelligence, and other applications of speech technologies with data compression based on the linear prediction model.