<p>The short-time magnitude spectrum (STFT-MS) captures the phonetic information and pitch of the speaker. The pitch effect can be normalized by generating an approximate excitation signal and deconvolving its response in the spectral domain. Building upon this motivation, this paper presents an approach for the normalization of the pitch in the spectral domain for the development of keyword spotting in out-of-domain pitch-mismatched test conditions. In the proposed approach, linear prediction (LP) analysis of the given speech signal is performed to generate the LP residual signal. Next, the Hilbert envelope (HE) of the LP residual signal is computed to enhance the glottal closure instants. The HE of the LP residual mostly captures the excitation source information. For each STFT-MS, the pitch effect is normalized by dividing the spectrum of the speech signal by the corresponding spectrum of the HE of the LP residual. The division of these two spectra is equivalent to the deconvolution of the excitation source from the response of the vocal tract filter. The Mel-frequency cepstral coefficient (MFCC) is then computed from the deconvolution spectra, termed excitation source normalized MFCC (ESN-MFCC). The pitch robustness of ESN-MFCC features is validated empirically by evaluating keyword spotting performance on a deep neural network-hidden Markov model-based KWS system under out-of-domain pitch-mismatched test conditions and comparing it with existing methods reported for the normalization of pitch and Wave2Vec-based self-supervised feature extractor. The experimental results presented in this study show that the ESN-MFCC feature surpasses all the explored features. For the zero-resource children’s KWS system, on the use of ESN-MFCC, the TWV improves to – 0.0051 and 0.0090 from the MFCC baseline of – 2.2189 and – 1.2135 for 10 and 20 keyword sets, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deconvolution of Source and System Responses for Development of KWS System in Out-of-Domain Test Conditions

  • Kaustav Das,
  • Gayadhar Pradhan

摘要

The short-time magnitude spectrum (STFT-MS) captures the phonetic information and pitch of the speaker. The pitch effect can be normalized by generating an approximate excitation signal and deconvolving its response in the spectral domain. Building upon this motivation, this paper presents an approach for the normalization of the pitch in the spectral domain for the development of keyword spotting in out-of-domain pitch-mismatched test conditions. In the proposed approach, linear prediction (LP) analysis of the given speech signal is performed to generate the LP residual signal. Next, the Hilbert envelope (HE) of the LP residual signal is computed to enhance the glottal closure instants. The HE of the LP residual mostly captures the excitation source information. For each STFT-MS, the pitch effect is normalized by dividing the spectrum of the speech signal by the corresponding spectrum of the HE of the LP residual. The division of these two spectra is equivalent to the deconvolution of the excitation source from the response of the vocal tract filter. The Mel-frequency cepstral coefficient (MFCC) is then computed from the deconvolution spectra, termed excitation source normalized MFCC (ESN-MFCC). The pitch robustness of ESN-MFCC features is validated empirically by evaluating keyword spotting performance on a deep neural network-hidden Markov model-based KWS system under out-of-domain pitch-mismatched test conditions and comparing it with existing methods reported for the normalization of pitch and Wave2Vec-based self-supervised feature extractor. The experimental results presented in this study show that the ESN-MFCC feature surpasses all the explored features. For the zero-resource children’s KWS system, on the use of ESN-MFCC, the TWV improves to – 0.0051 and 0.0090 from the MFCC baseline of – 2.2189 and – 1.2135 for 10 and 20 keyword sets, respectively.