In recent years, speech technology has received increasing attention from the scientific community, particularly when integrated into consumer products like home assistants and smartphones. However, it is important to emphasize that the efficiency of automatic speech recognition systems heavily depends on the preprocessing techniques used for the speech signal. Among these techniques, spectrogram resizing plays a crucial role in improving recognition accuracy. In this study, our objective is to investigate the impact of spectrogram resizing on the accuracy of a speech recognition system based on a convolutional neural network. We focused on the extraction and classification of an Amazigh isolated word corpus comprising 18 classes, which was collected from speakers in the Rif region of Morocco. The initial size of the Mel spectrogram was 96 × 96, and we explored resizing it to various dimensions. Additionally, we assessed the influence of different resizing interpolation techniques on speech recognition performance, including Bilinear, Lanczos, Gaussian, and others. Our findings indicate that resizing the Mel spectrogram to a resolution of 64 × 64 pixels using Bilinear interpolation yielded optimal performance, achieving a speech recognition accuracy of 93.66%. However, alternative interpolation techniques such as Lanczos 5 and area interpolation also produced viable options, offering reasonably high accuracy scores.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Impact of Spectrogram Resizing for Automatic Speech Recognition Case of Amazigh Isolated Word

  • Mohamed Daouad,
  • Fadoua Ataa Allah,
  • El Wardani Dadi

摘要

In recent years, speech technology has received increasing attention from the scientific community, particularly when integrated into consumer products like home assistants and smartphones. However, it is important to emphasize that the efficiency of automatic speech recognition systems heavily depends on the preprocessing techniques used for the speech signal. Among these techniques, spectrogram resizing plays a crucial role in improving recognition accuracy. In this study, our objective is to investigate the impact of spectrogram resizing on the accuracy of a speech recognition system based on a convolutional neural network. We focused on the extraction and classification of an Amazigh isolated word corpus comprising 18 classes, which was collected from speakers in the Rif region of Morocco. The initial size of the Mel spectrogram was 96 × 96, and we explored resizing it to various dimensions. Additionally, we assessed the influence of different resizing interpolation techniques on speech recognition performance, including Bilinear, Lanczos, Gaussian, and others. Our findings indicate that resizing the Mel spectrogram to a resolution of 64 × 64 pixels using Bilinear interpolation yielded optimal performance, achieving a speech recognition accuracy of 93.66%. However, alternative interpolation techniques such as Lanczos 5 and area interpolation also produced viable options, offering reasonably high accuracy scores.