Advancing Vietnamese Speech Emotion Recognition: The OrionNet Architecture and VESC Corpus
摘要
Speech Emotion Recognition (SER) has shown great potential for a wide range of applications; however, it remains a significant challenge for low-resource languages such as Vietnamese, primarily due to the severe scarcity of high-quality and naturalistic datasets. This study makes two main contributions to address this issue. First, we construct and introduce VESC (Vietnamese Emotional Speech Corpus), a Vietnamese emotional speech dataset comprising 904 audio samples (~72 m) collected from authentic sources such as films and television programs, involving 78 speakers. Second, we propose OrionNet, a lightweight 2D Convolutional Neural Network (CNN) architecture specifically designed to efficiently learn from multi-feature acoustic representations under limited-data conditions. Comprehensive experiments on the VESC dataset demonstrate that OrionNet achieves superior performance, attaining an accuracy of 95%, significantly outperforming state-of-the-art models. This study not only provides a valuable resource for the research community but also highlights the potential of specialized model design for tackling tasks in low-resource languages.