There are growing concerns regarding how easily accessible and rapidly advancing deep learning models can be used to create and synthesize hyper-realistic deepfake videos, which could potentially be misused for malicious purposes. In current trends, deep learning algorithms can create new faces, switch faces between two people in a film, change the gender of a subject, change facial expressions, and modify facial traits, and so on. Numerous fields may find applications for these formidable techniques for manipulating videos. However, if they are utilized for malicious intent like identity theft, phishing, or scams, they also threaten everyone. This study employs a combination of the Convolutional Neural Network (CNN) and Vision Transformer (ViT), referred to as the Convolutional Vision Transformer (CViT), to identify fake videos. In this approach, CNN is responsible for extracting useful features from the network’s hidden; layers, while the ViT utilizes these features and its attention mechanism to determine the authenticity of a video. The evaluation of this method is performed using the DeepFake Detection Challenge Dataset (DFDC).The proposed CViT model performed well in detecting the input videos into fake and original classes. The model obtained an average accuracy of 57.14%. The neural network models that were considered state-of-the-art were also outperformed by them.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fake Video Detection Using Deep Learning-Based Methods

  • Nishant Koshti,
  • V. Spoorthy,
  • Dheeraj K. Shringi,
  • Mayuri Popat

摘要

There are growing concerns regarding how easily accessible and rapidly advancing deep learning models can be used to create and synthesize hyper-realistic deepfake videos, which could potentially be misused for malicious purposes. In current trends, deep learning algorithms can create new faces, switch faces between two people in a film, change the gender of a subject, change facial expressions, and modify facial traits, and so on. Numerous fields may find applications for these formidable techniques for manipulating videos. However, if they are utilized for malicious intent like identity theft, phishing, or scams, they also threaten everyone. This study employs a combination of the Convolutional Neural Network (CNN) and Vision Transformer (ViT), referred to as the Convolutional Vision Transformer (CViT), to identify fake videos. In this approach, CNN is responsible for extracting useful features from the network’s hidden; layers, while the ViT utilizes these features and its attention mechanism to determine the authenticity of a video. The evaluation of this method is performed using the DeepFake Detection Challenge Dataset (DFDC).The proposed CViT model performed well in detecting the input videos into fake and original classes. The model obtained an average accuracy of 57.14%. The neural network models that were considered state-of-the-art were also outperformed by them.