The Whisper model, although advanced in speech recognition, suffers from slow speed and high memory usage. This study aims to address these issues while improving the efficiency of speech recognition without compromising accuracy. To achieve this, we combined Low-Rank adaptation and Tucker decomposition strategy. By specializing to the characteristics of different languages and helping the model better adapt to multilingual environments, we significantly enhanced the models’ trade-off relationships between accuracy vs. model size. Experimental results demonstrate that our approach achieves rapid and accurate speech transcription, making significant progress in multilingual environments and substantially reducing the Word Error Rate, providing an innovative solution for practical applications in multilingual speech processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Multilingual Speech Recognition with Tucker-Compressed Mixture of LoRAs

  • Ye Hong,
  • Guanghui Song,
  • Xitong Gao,
  • Tianhui Meng,
  • Juanjuan Zhao,
  • Kejiang Ye

摘要

The Whisper model, although advanced in speech recognition, suffers from slow speed and high memory usage. This study aims to address these issues while improving the efficiency of speech recognition without compromising accuracy. To achieve this, we combined Low-Rank adaptation and Tucker decomposition strategy. By specializing to the characteristics of different languages and helping the model better adapt to multilingual environments, we significantly enhanced the models’ trade-off relationships between accuracy vs. model size. Experimental results demonstrate that our approach achieves rapid and accurate speech transcription, making significant progress in multilingual environments and substantially reducing the Word Error Rate, providing an innovative solution for practical applications in multilingual speech processing.