Utilizing transfer learning based on the Wav2Vec2 pre-trained model has led to notable improvements in speech anti-spoofing capabilities. Nevertheless, the practical implementation of these models is restricted by the limitations of hardware and reduced performance when evaluating short utterances, which presents difficulties in incorporating them into streaming systems. In order to address these problems, we propose a novel training technique and a lightweight model for audio deepfake detection, explicitly tailored for short-duration utterances in resource-constrained environments. Our approach incorporates Voice Activity Detection (VAD) during both the training and inference stages, enhancing the accuracy of countermeasure models by focusing on speech rather than silence. We also introduce a knowledge distillation framework to compress state-of-the-art models efficiently. These methods significantly improve short utterances evaluation while maintaining robust performance for real-time applications. Experimental results demonstrate the effectiveness of our techniques in enhancing short utterance detection accuracy and ensuring feasibility in resource-limited scenarios, paving the way for more secure and reliable voice authentication systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Real-Time Audio Deepfake Detection in Resource-Limited Environments

  • Hung Dinh-Xuan,
  • Thien-Phuc Doan,
  • Sungkyu Han,
  • Kihun Hong,
  • Souhwan Jung

摘要

Utilizing transfer learning based on the Wav2Vec2 pre-trained model has led to notable improvements in speech anti-spoofing capabilities. Nevertheless, the practical implementation of these models is restricted by the limitations of hardware and reduced performance when evaluating short utterances, which presents difficulties in incorporating them into streaming systems. In order to address these problems, we propose a novel training technique and a lightweight model for audio deepfake detection, explicitly tailored for short-duration utterances in resource-constrained environments. Our approach incorporates Voice Activity Detection (VAD) during both the training and inference stages, enhancing the accuracy of countermeasure models by focusing on speech rather than silence. We also introduce a knowledge distillation framework to compress state-of-the-art models efficiently. These methods significantly improve short utterances evaluation while maintaining robust performance for real-time applications. Experimental results demonstrate the effectiveness of our techniques in enhancing short utterance detection accuracy and ensuring feasibility in resource-limited scenarios, paving the way for more secure and reliable voice authentication systems.