Improving Speaker Verification Back-End with Graph Neural Networks
摘要
Currently, research on speaker verification tasks is primarily concentrated on enhancing deep speaker models to extract high-quality speaker embeddings. Nevertheless, this speaker embeddings can be regarded as potential graph structures. Therefore, this paper introduces a graph neural network-based speaker verification back-end approach, treating the speaker embeddings derived from the front-end as graph structures and leveraging graph neural networks to uncover the relationships among embeddings, thereby achieving superior-quality speaker embeddings. In addition, we propose a group updating method to solve the problem of memory overflow when the number of nodes is excessive. Extensive experiments and ablation studies were carried out on the VoxCeleb dataset, and the experimental outcomes validate the effectiveness of our proposed graph neural network speaker back-end in significantly boosting the performance of speaker verification systems.