A Multi-aspect Multi-granularity Pronunciation Assessment Method Based on Branchformer Encoder and Hierarchical Aggregation
摘要
With the advancement of information technology, Computer-Assisted Pronunciation Training (CAPT) has become an effective method for non-native(L2) speakers to learn foreign language pronunciation. However, existing automatic pronunciation quality assessment methods have not fully leveraged the inter-granularity relationships and lack further extraction of contextual features at each granularity. To address these issues, this paper proposes Bfhaformer. Bfhaformer employs an LSTM-augmented BranchFormer encoder for encoding GOP features and reference phoneme features. Compared to Transformer encoders, the BranchFormer encoder introduces parallel branch structures, which enhances the capture of local features while retaining global feature information. Additionally, this paper aggregates features across different granularities within a hierarchical model structure. By aggregating and suprasegmental feature fusion of the encoded features at pronunciation granularity such as word level and utterance level, better attention is paid to local information at the current granularity and contextual hierarchical relationships. Experiments on the publicly available Speechocean762 dataset demonstrate that our proposed method significantly improves all metrics at all granularities compared to the baseline models.