Audio Splicing Forgery Detection via Discrepancy-Aware Meta-Learning
摘要
Amid rising audio forgeries, this paper introduces a zero-shot audio splicing detection framework based on discrepancy-aware meta-learning, enabling robust generalization across diverse manipulation types. The proposed method utilizes Constant-Q Transform (CQT) for feature extraction and employs discrepancy maps to detect subtle inconsistencies between authentic and forged audio. Through meta-learning, the framework captures manipulation-invariant features, achieving AUC scores of 94.19%, 95.47%, and 98.13% on TIMIT and ITKGP-SESC for splicing, pitch shifting, and time stretching, respectively. It also attained AP values of 91.53%, 93.11%, and 95.81%, with F1 scores of 90.49%, 92.68%, and 94.92%. These results highlight the model’s robustness in detecting forged audio. However, reliance on predefined manipulation types may limit adaptability to previously unseen attacks.