Who is Being Impersonated? Deepfake Audio Detection and Impersonated Identification via Extraction of Id-Specific Features
摘要
With the increasing concern about network security and model interpretability, fake audio detection is attracting significant attention. Although prior studies primarily concentrate on distinguishing authentic from fabricated audio, a vital yet underexplored dimension is the pinpointing within such detections, particularly in identifying the impersonated individual from counterfeit audio. In this paper, we propose a model capable of simultaneously detecting fake audio and identifying the speaker by extracting ID-specific features. Furthermore, we produce the Fake In-the-Wild Audio (FIA) dataset by expanding the “In-the-Wild” audio dataset. We adopt the advanced Text-to-Speech generation model, MetaVoice-1B, to generate fake audios based on the “In-the-Wild” audio dataset. We conducted a detailed analysis of MetaVoice-1B, focusing on its capabilities in generating realistic deepfake audio. Additionally, we identified its disadvantages and suggest potential future improvements. Experiments on the FIA dataset demonstrate the excellent performance of the proposed model, achieving a remarkable F1 score of 0.99 on the test set. Moreover, the model’s ability to output additional relevant information enhances overall cybersecurity by providing deeper insights into potentially fake content.