In the process of human-computer intelligent speech interaction, inferring the speaker's age from human speech signals presents a challenging task. The acoustic features related to age in speaker’s voice are complex, making it difficult for traditional machine learning methods to achieve comprehensive and accurate recognition results. This paper proposes a method for speaker age recognition based on a framework that integrates convolutional and self-attention mechanisms. Firstly, speech signals are transformed into spectrograms, and a CNN-Transformer Dual Branch Parallel Fusion Network (CTPF-Net) is designed to achieve a comprehensive extraction of global and local detail features of speech signals. Additionally, gender information is considered during training to perform unified age-gender recognition, achieving better accuracy than age recognition alone. Experiments and analysis on the Common Voice dataset demonstrate that the proposed model achieves an average accuracy of 84.5% in age recognition tasks. Moreover, without significantly increasing model complexity, the model can accurately differentiate speakers across different age segments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker Age Recognition Based on Convolution and Transformer Fusion Framework

  • Zheyan Zhang,
  • Renwei Li,
  • Kewei Chen

摘要

In the process of human-computer intelligent speech interaction, inferring the speaker's age from human speech signals presents a challenging task. The acoustic features related to age in speaker’s voice are complex, making it difficult for traditional machine learning methods to achieve comprehensive and accurate recognition results. This paper proposes a method for speaker age recognition based on a framework that integrates convolutional and self-attention mechanisms. Firstly, speech signals are transformed into spectrograms, and a CNN-Transformer Dual Branch Parallel Fusion Network (CTPF-Net) is designed to achieve a comprehensive extraction of global and local detail features of speech signals. Additionally, gender information is considered during training to perform unified age-gender recognition, achieving better accuracy than age recognition alone. Experiments and analysis on the Common Voice dataset demonstrate that the proposed model achieves an average accuracy of 84.5% in age recognition tasks. Moreover, without significantly increasing model complexity, the model can accurately differentiate speakers across different age segments.