Optimizing deepfake audio detection: fragment feature overlaid quantization model for high accuracy and efficiency
摘要
The increasing sophistication of deepfake technology necessitates the development of efficient methods to differentiate between Real and fake audio. This paper presents a robust model with Fragment Feature Overlaid, which integrates a feature fragmentation strategy to enhance model interpretability and robustness for audio authenticity detection. Integrating quantization techniques in neural network training and inference processes leads to significant enhancements in model efficiency without compromising accuracy. The proposed system conducts several quantization operations on the model to optimize it to permit accurate fake audio detection in a very short time. Subsequent experiments performed on an equal distribution of data prove that their quantized model sustains a high detection rate, with a comparatively lower execution burden for a practical intervention in the fight against audio deepfakes with possible use in media certification, investigative instruments, and safety in conversation. The analysis of the evaluation proves that Non-Uniform Quantization yields better results than other quantization methods. It reached a test error rate of 2.84%, a ROC-AUC of 0.9982, and a PR-AUC of 0.9981, which maintains high accuracy for detecting abnormalities and firmly establishes stable generalization performances.