Evaluating Suprasegmental Features for Phonological Fusion and Spectrogram-Based Speech Command Recognition
摘要
This study advances automatic speech recognition (ASR) by integrating multimodal information, combining textual transcripts with spatio-temporal features extracted from denoised audio signals, where median filtering is employed to reduce noise and preserve key structural patterns. It investigates spectrogram analysis and phonology for enhancing recognition accuracy of speech commands. A two-pronged strategy is utilized, applying the Speech2Text transformer for accurate transcript extraction and the Swin transformer for effective spectrogram processing. The research confirms its approach on the Google Speech Command dataset (version 2) (GSCD v2) with both 10- and 35-speech command classes. Mel spectrograms are computed at 256 × 256 resolution and classified with ImageNet and the Tiny Swin Transformer v2. The work also incorporates a grapheme-to-phoneme (G2P) model to convert textual transcripts into phoneme sequences. These phonemes are then segmented through phoneme slicing to extract key linguistic features. Special attention is given to different phoneme classes—fricatives, nasals, plosives, glides, liquids, traps, trills, lateral fricatives and vowels—by considering their distinct articulatory characteristics. The research utilizes ablation analysis to measure the effect of spectrograms and phonological features on ASR performance. A late fusion strategy is employed to integrate the probabilistic scores derived from both phoneme-based and spectrogram-based techniques. This fusion significantly enhances the accuracy of the ASR system. The proposed model achieves an impressive 99.87% accuracy on a 10-word classification task and 98.53% on a 35-word classification task, outperforming existing state-of-the-art benchmarks. Through the fusion of linguistic knowledge and deep learning, the research emphasizes the pivotal role played by phonological analysis in speech processing. The results not only enhance ASR accuracy but also establish a new benchmark in speech recognition, illustrating the strength of multimodal methods in pushing the boundaries of ASR technology.