Accent Recognition with Auxiliary Task and Contrastive Learning
摘要
We propose a multi-task learning approach that integrates Automatic Speech Recognition (ASR) and Accent Recognition (AR) tasks. By utilizing a shared ASR encoder alongside a Transformer-based AR decoder, our model enhances the extraction of relevant accent features, overcoming challenges posed by non-joint frameworks. With contrastive learning, geographical information is effectively leveraged. Experimental results validate the effectiveness of our proposed method, demonstrating improved accent classification performance and indicate the features important for AR task.