AO-UAP: An Adaptive Universal Adversarial Perturbation Generation for Speech Recognition Models
摘要
Speech recognition (SR) has become an essential component in multiple applications, such as intelligent voice control systems and smart voice assistants, whose robustness is crucial for large-scale deployment. To explore the vulnerability of SR models, one classic approach crafts adversarial examples for each audio sample, which is audio-dependent and lacks practicality. Recent researches attempt to craft a single perturbation applied to most samples. However, they generate a perturbation for an utterance at once, leading to slow convergence. To tackle these challenges, in this paper we propose AO-UAP, a novel approach to adaptively generate audio universal adversarial perturbations (UAPs) for untargeted adversarial attack tasks. We utilize a penalty-based generation algorithm and leverage preserved momentum and gradients to optimize UAP with individualized learning rates, contributing to fast convergence. Our approach boosts the performance of baseline UAP-HC by a 2.14% fooling rate against X-vector models. The comprehensive experiments on two representative model architectures demonstrate the effectiveness and efficiency of our method, reaching 88.92% and 87.28% fooling rates on Speech Commands and AudioMNIST test sets, respectively.