Seeing Through the Lies: A Vision Transformer-Based Solution
摘要
Deepfake technology has become increasingly sophisticated and poses a growing threat to society, as it can be used to create convincing fake videos for malicious purposes. Therefore, detecting deepfakes has become crucial to ensure the authenticity of visual content. This paper proposes a novel deepfake detection approach that utilizes InceptionResNetV2 for feature extraction and Vision Transformer (ViT) with Nyström Attention mechanism for classification. Our model achieves high accuracy, AUC, precision, and recall on the CelebDFv2 dataset and outperforms state-of-the-art techniques like EfficientNet, XceptionNet, and ResNet. We also perform evaluations on the CelebDFv1 and DFDC datasets. Furthermore, we conduct a comprehensive literature survey to highlight the existing research in deepfake detection. We provide a roadmap for future work in this field, highlighting potential research directions that can further improve the effectiveness and applicability of deepfake detection models. The proposed model shows great potential for real-world applications, and the suggested future work can help to address some of the remaining challenges in this area.