Resformer: Local Frame-Level Feature and Global Segment-Level Feature Joint Learning for Speaker Verification
摘要
In this paper, we propose a hybrid network structure to achieve more discriminant feature representations for speaker recognition, termed Resformer, local frame-level features are extracted by convolution operation, and global segment-level features are extracted by transformer, which the information learned from the frame-level features is used to guide the segment-level feature learning, to take advantage of convolutional operations and self-attention mechanisms for enhanced representation learning. Resformer uses an interactive approach to combine local features and global representations at varying resolutions. With a concurrent structure, Resformer ensures that both local features and global representations are preserved to the fullest extent possible. Extensive experiments are performed on VoxCeleb, a public SV dataset, and the findings indicate that Resformer outperforms existing deep embedding architectures by a margin of at least 10–37% on the test set. Our proposed approaches are shown to achieve significant improvement over prior methods through the ablation experiments.