An Intonation-Based Black-Box Generative Adversarial Attack Method for Audio
摘要
Deep learning has achieved remarkable success in the field of automatic speech recognition (ASR) and has been widely applied in real-world scenarios. However, ASR models based on deep neural networks exhibit vulnerabilities against adversarial examples. Current research on adversarial ASR models predominantly focuses on white-box methods, whereas real-world environments typically involve black-box settings. Black-box attacks face challenges such as low success rates and high detectability. In particular, previous genetic algorithm-based methods require extensive query budgets, significantly increasing the risk of detection. To address these issues, this paper innovatively proposes a black-box adversarial audio attack method based on intonation modulation. Specifically, an adaptive particle swarm optimization algorithm is employed to fine-tune intonation parameters precisely, while the original audio and its corresponding transcription are incorporated into the generative model to craft adversarial examples. Experimental evaluations on DeepSpeech, CMU Sphinx, and iFlytek ASR systems demonstrate that the proposed method achieves a high attack success rate while effectively reducing query overhead and preserving perceptual audio quality.