The proposed research presents a comprehensive method for diagnosing and detecting mispronunciations, using a Variational Autoencoder (VAE) to improve performance. The proposed approach combines several feature variables and modalities for a more efficient analysis. Using VAE as an audio encoder, audio data representations are captured to learn a compact and useful latent space representation. Bi-directional Long Short-Term Memory (Bi-LSTM) network is used as a phoneme encoder to extract phonetic information from the input data. Transformer-based decoder is incorporated to decode the learnt representations and provide sequences associated with speech pattern pronunciation. To improve the model’s comprehension of the speaker’s pronunciation material, Multi-Modal Fusion incorporates phoneme information both before and during the different phases of VAE. The implementation of a secondary decoding mechanism is part of the decoding process. This method entails returning the decoded sequence to the decoder for a further round of decoding. This improves mispronunciation diagnosis and identification by overcoming the difficulty of incomplete prior knowledge during the initial decoding stage. The L2 Arctic English dataset is used for the experimental evaluation of the model. Significant improvements are shown as compared to the baseline model which used Squeezeformer as an encoder (Guo et al. in “Multi-feature and multi-modal mispronunciation detection and diagnosis method based on the squeezeformer encoder”. IEEE Access (99):1–1). The accuracy increased slightly from 96.19 to 96.27% when compared to the baseline model. Moreover, when contrasted with various models, there was an average accuracy improvement of 3.2%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VAE MDD: VAE-Based Multi-modal analysis for Enhanced Mispronunciation Recognition

  • Shreya Sharma,
  • Chesta Agarwal,
  • Pinaki Chakraborty

摘要

The proposed research presents a comprehensive method for diagnosing and detecting mispronunciations, using a Variational Autoencoder (VAE) to improve performance. The proposed approach combines several feature variables and modalities for a more efficient analysis. Using VAE as an audio encoder, audio data representations are captured to learn a compact and useful latent space representation. Bi-directional Long Short-Term Memory (Bi-LSTM) network is used as a phoneme encoder to extract phonetic information from the input data. Transformer-based decoder is incorporated to decode the learnt representations and provide sequences associated with speech pattern pronunciation. To improve the model’s comprehension of the speaker’s pronunciation material, Multi-Modal Fusion incorporates phoneme information both before and during the different phases of VAE. The implementation of a secondary decoding mechanism is part of the decoding process. This method entails returning the decoded sequence to the decoder for a further round of decoding. This improves mispronunciation diagnosis and identification by overcoming the difficulty of incomplete prior knowledge during the initial decoding stage. The L2 Arctic English dataset is used for the experimental evaluation of the model. Significant improvements are shown as compared to the baseline model which used Squeezeformer as an encoder (Guo et al. in “Multi-feature and multi-modal mispronunciation detection and diagnosis method based on the squeezeformer encoder”. IEEE Access (99):1–1). The accuracy increased slightly from 96.19 to 96.27% when compared to the baseline model. Moreover, when contrasted with various models, there was an average accuracy improvement of 3.2%.