MSCACodec: A Low-Rate Neural Speech Codec With Multi-scale Residual Channel Attention
摘要
In the development of modern communication technology, although wideband speech coding provides high-fidelity speech transmission, its high bandwidth requirement limits its application in resource-constrained environments. Thus, narrowband speech coding is still of great significance. Recently, end-to-end neural speech coding has made significant progress and demonstrated superior compression performance over traditional methods. However, existing methods are limited in reconstructing details, especially in low birate environments. To address this, we introduce MSCACodec, a narrowband-based neural speech codec that achieves advanced performance at low bitrates. MSCACodec adopts a multi-scale residual and channel attention feature fusion method to selectively focus on multi-scale information to enhance feature representation, solving the problem of inconsistent hierarchical information caused by multi-scale feature fusion. In addition, we also propose a Temporal Convolutional Gated Recurrent Unit (TCGRU) module, which combines temporal convolutional networks and gated recurrent units to enhance the reconstruction quality using global context and gating mechanisms. The experimental results show that, whether in subjective or objective evaluation, MSCACodec achieves higher quality reconstructed speech than Encodec and HiFiCodec at bitrates of 1.2 kbps and 2.4 kbps, and is even better than LyraV2 and Opus at 6 kbps.