<p>The advent of synthetic voice generation technology, it has created transformative capabilities and brought about significant challenges, especially in voice authentication, digital forensics, and security. Deepfake audio is one type of deepfake content that is artificially created or altered to sound like a human voice and can be highly dangerous through identity theft, misinformation, and exposure of confidentiality. True identification of such digitally imitated voices is critical in ensuring that we are curtailing these risks and ensuring the sanctity of the voice-based systems. The study presents a systematic process of Digitally Generated Voice Detection using Machine Learning and Deep Learning systems. The data were properly prepared in a manner that separates the real and fake audio files, representing the extracted features in isolation and using audio files to train Random Forest, Deep Neural Network (DNN) model, XGBoost, and a combination of DNN + XGBoost model. Characteristics such as HNR, pitch variance, frequency range, intensity, mean pitch, chroma, mel spectrogram, and spectral contrast are extracted from the signals. MFCCs, chroma, mel, and spectral contrast are relatively common in speech and audio processing because they represent the important features of an audio signal, including how the human auditory system responds to sound. This extraction process gives the system the capability to notice minute variations between human and artificial voices, thus playing an essential role in the detection of the voice characteristics. These results indicate that the combination of DNN and XGBoost model could perform the best in terms of accuracy, in terms of precision, and computation. This research addresses the increasing threats that deepfake audio represents, thus advancing the security frameworks, preventing the potential harm that this type of voice fraud can cause, and developing the acoustic analysis of sound. The proposed system will offer an easy-to-scale, advanced technology to differentiate between the human and artificial voice and potentially form the basis of more advanced uses of deepfakes detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distinguishing Synthetic from Human Speech: A Machine Learning and Deep Learning Approach

  • Ashok Khedkar,
  • Ruta Chaudhari,
  • Richa Rathi,
  • Saee Gade

摘要

The advent of synthetic voice generation technology, it has created transformative capabilities and brought about significant challenges, especially in voice authentication, digital forensics, and security. Deepfake audio is one type of deepfake content that is artificially created or altered to sound like a human voice and can be highly dangerous through identity theft, misinformation, and exposure of confidentiality. True identification of such digitally imitated voices is critical in ensuring that we are curtailing these risks and ensuring the sanctity of the voice-based systems. The study presents a systematic process of Digitally Generated Voice Detection using Machine Learning and Deep Learning systems. The data were properly prepared in a manner that separates the real and fake audio files, representing the extracted features in isolation and using audio files to train Random Forest, Deep Neural Network (DNN) model, XGBoost, and a combination of DNN + XGBoost model. Characteristics such as HNR, pitch variance, frequency range, intensity, mean pitch, chroma, mel spectrogram, and spectral contrast are extracted from the signals. MFCCs, chroma, mel, and spectral contrast are relatively common in speech and audio processing because they represent the important features of an audio signal, including how the human auditory system responds to sound. This extraction process gives the system the capability to notice minute variations between human and artificial voices, thus playing an essential role in the detection of the voice characteristics. These results indicate that the combination of DNN and XGBoost model could perform the best in terms of accuracy, in terms of precision, and computation. This research addresses the increasing threats that deepfake audio represents, thus advancing the security frameworks, preventing the potential harm that this type of voice fraud can cause, and developing the acoustic analysis of sound. The proposed system will offer an easy-to-scale, advanced technology to differentiate between the human and artificial voice and potentially form the basis of more advanced uses of deepfakes detection.