Dhivehi Speech Recognition: A Multimodal Approach for Dhivehi Language in Resource-Constrained Settings
摘要
This study addresses the challenging research problem of effectively integrating multiple data modalities, such as imperfect text transcripts and raw audio, within a deep learning framework for Dhivehi speech recognition. The focus is on resource-constrained settings, where limited prior research has been conducted. Our proposed approach employs a multimodal strategy, integrating audio, text, and image features through a late fusion technique. Specifically, we combine audio-to-spectrogram processed and passed to ResNet, phone embeddings, and 13 mel frequency cepstral coefficients (MFCC) to recognize Dhivehi speech spoken in Malaysia. To extract phonemic information, we utilize Carnegie Mellon University (CMU) phonemes obtained from text transcripts generated using the Speech2Text transformer on OPUS audio files. For phoneme embeddings, we leverage Language-Agnostic Sentence Representations (LASER), feeding these embeddings into feedforward neural networks. Our approach synergizes the strengths of audio, text, and image modalities to enhance recognition accuracy. In our experiments, we focus on the Dhivehi language within the extensive Multilingual Spoken Word Corpus (MSWC) dataset. By fusing posterior class probabilities from deep bidirectional long short-term memory (deep-BiLSTM) and feedforward neural models trained on individual modalities, we achieve an impressive accuracy of 96.42% when tested on 11 subjects within the MSWC dataset. Extensive testing and thorough ablation analyses demonstrate that our proposed method outperforms existing state-of-the-art systems. This superior performance can be attributed to the exceptional classification accuracies achieved by our approach. In conclusion, our study presents a robust and innovative multimodal approach that effectively combines audio, text, and image modalities for Dhivehi speech recognition in resource-constrained settings, showcasing superior accuracy compared to existing methodologies.