Deep learning has achieved remarkable success in the field of automatic speech recognition (ASR) and has been widely applied in real-world scenarios. However, ASR models based on deep neural networks exhibit vulnerabilities against adversarial examples. Current research on adversarial ASR models predominantly focuses on white-box methods, whereas real-world environments typically involve black-box settings. Black-box attacks face challenges such as low success rates and high detectability. In particular, previous genetic algorithm-based methods require extensive query budgets, significantly increasing the risk of detection. To address these issues, this paper innovatively proposes a black-box adversarial audio attack method based on intonation modulation. Specifically, an adaptive particle swarm optimization algorithm is employed to fine-tune intonation parameters precisely, while the original audio and its corresponding transcription are incorporated into the generative model to craft adversarial examples. Experimental evaluations on DeepSpeech, CMU Sphinx, and iFlytek ASR systems demonstrate that the proposed method achieves a high attack success rate while effectively reducing query overhead and preserving perceptual audio quality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Intonation-Based Black-Box Generative Adversarial Attack Method for Audio

  • Wen Cui,
  • Pengchuan Wang,
  • Qianmu Li

摘要

Deep learning has achieved remarkable success in the field of automatic speech recognition (ASR) and has been widely applied in real-world scenarios. However, ASR models based on deep neural networks exhibit vulnerabilities against adversarial examples. Current research on adversarial ASR models predominantly focuses on white-box methods, whereas real-world environments typically involve black-box settings. Black-box attacks face challenges such as low success rates and high detectability. In particular, previous genetic algorithm-based methods require extensive query budgets, significantly increasing the risk of detection. To address these issues, this paper innovatively proposes a black-box adversarial audio attack method based on intonation modulation. Specifically, an adaptive particle swarm optimization algorithm is employed to fine-tune intonation parameters precisely, while the original audio and its corresponding transcription are incorporated into the generative model to craft adversarial examples. Experimental evaluations on DeepSpeech, CMU Sphinx, and iFlytek ASR systems demonstrate that the proposed method achieves a high attack success rate while effectively reducing query overhead and preserving perceptual audio quality.