Speaker recognition with global information modelling of raw waveforms
摘要
In recent years, methods to extract speaker-embedding information directly from raw waveforms have received much attention, with good results achieved by the RawNet3 network. However, the RawNet3 model only uses convolutional neural networks (CNNs) to extract speaker features directly from the raw waveform, which limits the perceptual field and leads to the model’s inability to learn more speaker features with long-term dependencies. This paper proposes a novel speaker recognition model with global information modelling of raw waveform, called GIMR-Net, which is able to extract more speaker features with long-term dependencies. The model uses the transformer structure in the network to extract global information and combines it with the CNN structure to extract local information, thus achieving the ability to model the global information of the raw waveform. Experiments show that the proposed GIMR-Net model is effective and outperforms the RawNet3 model on the Free ST Chinese Mandarin Corpus dataset. Specifically, the equal error rate of GIMR-Net is 1.22, a 12.9