In recent years, the rapid development of voice synthesis technologies has led to an increasing concern about the abuse of fake human voices for malicious purposes, such as deepfake audio, spam calls and social engineering attacks. This paper proposes a novel deep learning-based model to effectively identify counterfeit human voices generated by various voice synthesis algorithms. The proposed model employs a combination of Dense-Style Network to capture both spectral and temporal features of human speech. The model is extensively evaluated on ASVspoof 2019 datasets. The experimental results indicate that our model achieves competitive performance compared to existing methods and has a certain degree of anti-compression ability. In addition, anti-compression research was conducted to investigate the recognition performance of the model in response to compressed speech. Our findings pave the way for further research in combating against the misuse of artificially generated human voices and sound authenticity verification in general.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detection of Speech Spoofing Based on Dense Convolutional Network

  • Yong Wang,
  • Xiaozong Chen,
  • Yifang Chen,
  • Shunsi Zhang

摘要

In recent years, the rapid development of voice synthesis technologies has led to an increasing concern about the abuse of fake human voices for malicious purposes, such as deepfake audio, spam calls and social engineering attacks. This paper proposes a novel deep learning-based model to effectively identify counterfeit human voices generated by various voice synthesis algorithms. The proposed model employs a combination of Dense-Style Network to capture both spectral and temporal features of human speech. The model is extensively evaluated on ASVspoof 2019 datasets. The experimental results indicate that our model achieves competitive performance compared to existing methods and has a certain degree of anti-compression ability. In addition, anti-compression research was conducted to investigate the recognition performance of the model in response to compressed speech. Our findings pave the way for further research in combating against the misuse of artificially generated human voices and sound authenticity verification in general.