The asynchronous real-time universal adversarial perturbation generation method for practical speaker recognition systems
摘要
Current adversarial attacks for speaker verification systems suffer from the following limitations: Existing methods typically require attackers to generate sample-specific perturbations tailored to individual audio clips to accommodate diverse attack scenarios, resulting in poor universality. Moreover, the attack process strictly requires precise temporal synchronization between the audio signal and perturbations, making real-time attack implementations particularly challenging. To address these challenges, this paper proposes the asynchronous real-time untargeted universal perturbation attack method for streaming audio, which effectively disrupts systems by playing perturbations at arbitrary timestamps. We design a novel optimization framework combining class-wise C&W loss with PGD methods. By jointly optimizing decision boundary perturbations across multiple user utterances, this addresses generalizability limitations, enabling single perturbations to effectively attack different speech samples from the same speaker. Then, we innovatively introduce a temporal delay domain adversarial training mechanism that dynamically injects randomized time shift parameters during optimization. This breaks the strict input synchronization constraints of conventional methods, enabling asynchronous attacks triggered arbitrarily within permissible time windows. Incorporating Room Impulse Response simulations during perturbation generation, we enhance the robustness of audio adversarial perturbations against real-world propagation losses. This ensures effectiveness even after airborne transmission in physical environments. Detailed experimental results demonstrate that our attack achieves a success rate of 95.10% in digital attacks and 100% success rate in physical environments.