Detection of Fake Audio: A Deep Learning-Based Comprehensive Survey
摘要
Deepfake audio is systematically generated audio used for defaming, blackmailing, and other bad purposes. However, audio deepfakes can also be used for other good purposes like for film shoots to reshoot the scene with different audios that will help to save a lot of cost. Deep learning is used for manipulating audio deepfakes so that humans cannot differentiate the audio. With human senses, it is impossible to detect audio deepfake, but with the enhancement in technologies, it is getting much more difficult to detect audio deepfakes. This paper will provide a survey of various methods used for detecting deep fake audio and a wide discussion on trends, challenges, and remarks for different methodologies used. We also figure out measurements for gaining unique features of human speech Methods surveyed in this paper are comprehensive and it will provide value to get an overview of deepfake detection methods.