CloneAI: A Deep Learning-Based Approach for Cloned Voice Detection
摘要
Voice cloning technology, which involves the creation of synthetic voices that imitate those of real individuals, has become increasingly sophisticated. To combat this menace, researchers have developed various machine learning and signal processing-based techniques for detecting fake voices. This research presents CloneAI, a convolutional neural network-based fake voice detector having four convolutional layers that uses Linear Frequency Cepstral Coefficients features and Mel spectrogram to analyse various speech features. The proposed detector’s efficacy was measured against a previously unseen dataset containing both synthetic and natural voices, and its generalisability was examined with manually recorded samples of natural speech. The model gave promising results, with an accuracy of 99.99% on testing data and all correct classification on manually recorded unseen real-world data. The results demonstrate the effectiveness of the proposed approach in detecting fake voices with high accuracy and suggest its potential for real-world applications in security, law enforcement, and other domains.