Defend from Scratch: A Diffusion-Based Proactive Defense Method for Unauthorized Speech Synthesis
摘要
With the advent of Deep Learning (DL), speech synthesis technologies have made remarkable progress, enabling the creation of highly realistic human voices. Although this technology offers numerous benefits, it also introduces substantial security risks, particularly through “Deepfake” speech attacks. These attacks pose severe threats to personal security and societal trust. Existing defense mechanisms, primarily focused on post-attack detection, fall short when confronted with the advanced sophisticate speech synthesis techniques. Worse still, these defenses are often insufficient because significant harm may have already been inflicted by the time deepfake audio is identified. In this paper, we present a novel proactive defense method against unauthorized speech synthesis named “Defend from Scratch” (DFS). By leveraging the idea of adversarial examples, our method can proactively hinder the creation of deepfake speeches. To do that, we propose to incorporate a pretrained decoupled denoising diffusion model (DDDM) to introduce robust and imperceptible adversarial perturbations into the source audio, which enables effectively defense against adaptive attacks without significant audio quality downgrade. Furthermore, Enhanced defenses have been implemented through the application of ensemble learning, which extend its capability to counter a diverse range of threats, thereby ensuring robust voice privacy protection without compromising the integrity or usability of the original audio. Experiments show that our approach stands out for its efficacy in maintaining a high standard of voice privacy in the face of emerging Deepfake technologies.