Enhancing Deepfake Detection Accuracy with Vision Transformers: A Generative AI Approach
摘要
Deepfake images are increasingly impacting everyday life, posing significant challenges to society. As the technology behind deepfakes evolves and becomes more accessible, various categories of deepfake photos emerge. Concurrently, deepfake detection methods are improving, ranging from basic feature analysis to deep learning approaches. However, there is no consistent method that can completely detect such images. The objective of this research is to provide an overview of current deepfake detection techniques and to examine the accuracy of Vision Transformer (ViT) based models when analyzing and detecting deep fake images. We implement a ViT model-based deepfake detection technique that is trained and tested on a hybrid dataset of FaceForensics++, Celeb-DF, and DFDC, which include both real and fake content images, the proposed ViT model achieves accuracy rates of 94.3%, 92.5%, and 95.1% against FaceForensics++, Celeb-DF, and DFDC, respectively. These results demonstrate the robustness of the ViT model's performance and ability to combat deepfake propagation.