AudioFormer: Channel Audio Encoder Based on Multi-granularity Features
摘要
To solve the problem of poor standardized feature extraction methods for speech emotion recognition tasks and insufficient depth representation capability for extracting acoustic samples, we first propose a Multi-granularity feature extraction method that takes into account the integrity of data features and overcomes the redundancy of existing feature extraction methods; secondly, we propose a Channel Audio Encoder Model that uses different Feature Encoders to extract High-order features. Experiments show that the proposed Multi-granularity feature-based Channel Audio Encoder achieves state-of-the-art performance in the IEMOCAP dataset. The method also experiments on a real-scene dataset to demonstrate its usability and provide a reference for aiding the diagnosis of mental illness.