The rapid progress of deep learning techniques has led to the generation of very realistic fake content by superimposing or replacing existing images, videos, or audio with highly realistic alternative content. These manipulations often involve the faces or voices of individuals, creating convincing but entirely fabricated representations. Due to the potential misuse of deepfakes for malicious purposes, the development of detection techniques and policies to mitigate the harmful effects of deepfakes has become an important area of research and societal concern. In this paper, we combine both spatial and frequency features to develop a simple yet effective model to detect deepfakes. Specifically, we use a one-level wavelet transform to decompose the grayscale image of an input color image in four subbands. We then rescale the horizontal, vertical, and diagonal subband images to the original resolution to capture the facial manipulations along three major orientations. Finally, we expand the input color images by appending the rescaled wavelet subbands to obtain an input image of six channels: the first three channels capture the spatial features and the next three channels capture the frequency features. These expanded input images are fed into the VGG19 backbone to detect deepfakes with an improved detection performance. We perform both within and cross domain evaluations to compare the performance of the proposed model and state-of-the-art peer models in terms of Area Under Curve (AUC) and Equal Error Rate (EER) metrics. Our extensive experimental results demonstrate that the proposed wavelet-based VGG19 model offers a more robust solution than the peer wavelet-based Xception model and both VGG19 and Xception baseline models to combating the proliferation of fake multimedia content on digital platforms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detection of Deepfakes Using Wavelet-Based Convolutional Neural Network

  • Supriyo Sadhya,
  • Xiaojun Qi

摘要

The rapid progress of deep learning techniques has led to the generation of very realistic fake content by superimposing or replacing existing images, videos, or audio with highly realistic alternative content. These manipulations often involve the faces or voices of individuals, creating convincing but entirely fabricated representations. Due to the potential misuse of deepfakes for malicious purposes, the development of detection techniques and policies to mitigate the harmful effects of deepfakes has become an important area of research and societal concern. In this paper, we combine both spatial and frequency features to develop a simple yet effective model to detect deepfakes. Specifically, we use a one-level wavelet transform to decompose the grayscale image of an input color image in four subbands. We then rescale the horizontal, vertical, and diagonal subband images to the original resolution to capture the facial manipulations along three major orientations. Finally, we expand the input color images by appending the rescaled wavelet subbands to obtain an input image of six channels: the first three channels capture the spatial features and the next three channels capture the frequency features. These expanded input images are fed into the VGG19 backbone to detect deepfakes with an improved detection performance. We perform both within and cross domain evaluations to compare the performance of the proposed model and state-of-the-art peer models in terms of Area Under Curve (AUC) and Equal Error Rate (EER) metrics. Our extensive experimental results demonstrate that the proposed wavelet-based VGG19 model offers a more robust solution than the peer wavelet-based Xception model and both VGG19 and Xception baseline models to combating the proliferation of fake multimedia content on digital platforms.