Speaker Age Recognition Based on Convolution and Transformer Fusion Framework
摘要
In the process of human-computer intelligent speech interaction, inferring the speaker's age from human speech signals presents a challenging task. The acoustic features related to age in speaker’s voice are complex, making it difficult for traditional machine learning methods to achieve comprehensive and accurate recognition results. This paper proposes a method for speaker age recognition based on a framework that integrates convolutional and self-attention mechanisms. Firstly, speech signals are transformed into spectrograms, and a CNN-Transformer Dual Branch Parallel Fusion Network (CTPF-Net) is designed to achieve a comprehensive extraction of global and local detail features of speech signals. Additionally, gender information is considered during training to perform unified age-gender recognition, achieving better accuracy than age recognition alone. Experiments and analysis on the Common Voice dataset demonstrate that the proposed model achieves an average accuracy of 84.5% in age recognition tasks. Moreover, without significantly increasing model complexity, the model can accurately differentiate speakers across different age segments.