<p>The ability to understand and respond to speech emotions and speaker identities can enhance human-robot interaction and enable more intuitive control of smart home devices, creating more personalized and responsive environments. Wav2vec 2.0 models excel in speech emotion recognition (SER) and speaker recognition (SSR) but are computationally demanding, hindering deployment on resource-limited devices. This paper addresses this challenge by introducing a novel multi-step compression methodology designed to reduce model size while maintaining high accuracy in both tasks. Our approach optimizes a pretrained wav2vec 2.0 model through a combination of encoder layer pruning, fine-tuning through knowledge distillation with a specially designed loss function that leverages both multitask labeled and unlabeled data, a novel neuron merging algorithm, and quantization. Experiments on the RAVDESS dataset provided valuable insights for tuning the proposed approach, demonstrating that our compressed models achieve a 9×–17× reduction in size and a 3.87×–8.57× acceleration on PC-CPU while attaining SER accuracy up to 95.74% and SSR accuracy up to 99.72%. These insights guided our subsequent large-scale evaluation on the IEMOCAP dataset using a Raspberry Pi 4, where our compressed models achieved a size reduction of 9.85×–10.26× and an acceleration of 3.62×–3.84× while maintaining high accuracies in SER, reaching 91.10%, and SSR, reaching 98.96%. This innovative compression method enables the deployment of high-performance speech recognition models on resource-constrained devices, facilitating their integration into smart home systems and robotics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient compression of wav2vec 2.0 for edge deployment in speech emotion & speaker recognition

  • Abderrahmane Bendahmane,
  • Rafe Alasem

摘要

The ability to understand and respond to speech emotions and speaker identities can enhance human-robot interaction and enable more intuitive control of smart home devices, creating more personalized and responsive environments. Wav2vec 2.0 models excel in speech emotion recognition (SER) and speaker recognition (SSR) but are computationally demanding, hindering deployment on resource-limited devices. This paper addresses this challenge by introducing a novel multi-step compression methodology designed to reduce model size while maintaining high accuracy in both tasks. Our approach optimizes a pretrained wav2vec 2.0 model through a combination of encoder layer pruning, fine-tuning through knowledge distillation with a specially designed loss function that leverages both multitask labeled and unlabeled data, a novel neuron merging algorithm, and quantization. Experiments on the RAVDESS dataset provided valuable insights for tuning the proposed approach, demonstrating that our compressed models achieve a 9×–17× reduction in size and a 3.87×–8.57× acceleration on PC-CPU while attaining SER accuracy up to 95.74% and SSR accuracy up to 99.72%. These insights guided our subsequent large-scale evaluation on the IEMOCAP dataset using a Raspberry Pi 4, where our compressed models achieved a size reduction of 9.85×–10.26× and an acceleration of 3.62×–3.84× while maintaining high accuracies in SER, reaching 91.10%, and SSR, reaching 98.96%. This innovative compression method enables the deployment of high-performance speech recognition models on resource-constrained devices, facilitating their integration into smart home systems and robotics.