The traditional voiceprint collection method usually requires manual construction of a voiceprint library, and a specific person is responsible for establishing and sorting it, which is cumbersome and time-consuming, which limits the application efficiency of voiceprint recognition technology and is difficult to meet the current diversified application needs. In this paper, an automatic voiceprint acquisition method based on deep learning is proposed. In this method, the speaker separation algorithm is used to automatically segment the dialogue audio into audio segments of different speakers, and then the speech recognition technology is used to convert these fragments into text, and finally the speaker's name is extracted from the text through a large language model, and matched with the real name in the user list to realize the automatic collection of voiceprints. This method effectively simplifies the voiceprint collection process, and automatically stores the voiceprint into the voiceprint library without manual intervention, which greatly improves the collection efficiency. This technology automatically collects the voiceprints of different speakers directly in the application environment, providing more robust recognition performance for the voiceprint recognition model, and is suitable for security verification, identity recognition and other fields, and has a wide range of application prospects.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Automatic Voiceprint Acquisition Method Based on Deep Learning

  • Quan Sun,
  • Ronghua Zhang,
  • Yingchun Liu,
  • Zemeng Liu,
  • Wei Chen

摘要

The traditional voiceprint collection method usually requires manual construction of a voiceprint library, and a specific person is responsible for establishing and sorting it, which is cumbersome and time-consuming, which limits the application efficiency of voiceprint recognition technology and is difficult to meet the current diversified application needs. In this paper, an automatic voiceprint acquisition method based on deep learning is proposed. In this method, the speaker separation algorithm is used to automatically segment the dialogue audio into audio segments of different speakers, and then the speech recognition technology is used to convert these fragments into text, and finally the speaker's name is extracted from the text through a large language model, and matched with the real name in the user list to realize the automatic collection of voiceprints. This method effectively simplifies the voiceprint collection process, and automatically stores the voiceprint into the voiceprint library without manual intervention, which greatly improves the collection efficiency. This technology automatically collects the voiceprints of different speakers directly in the application environment, providing more robust recognition performance for the voiceprint recognition model, and is suitable for security verification, identity recognition and other fields, and has a wide range of application prospects.