Distinction Between AI-Generated Images and Actual Images Using Swin-Transformer
摘要
The rapid advancement of generative AI has enabled the creation of highly realistic synthetic images and videos, raising concerns about privacy invasion, misinformation, and malicious misuse. These risks pose significant legal, ethical, and security challenges, necessitating robust detection mechanisms to prevent harm. Existing detection methods struggle with generalization and adversarial robustness, particularly when identifying high-quality, expertly manipulated synthetic content. Many prior studies rely on handcrafted feature extraction or limited dataset diversity, reducing their applicability in real-world scenarios. To address these limitations, this study introduces a Swin-Transformer-based detection model that enhances robustness by capturing both local and global visual features, improving accuracy across various synthetic media types. Unlike conventional CNN and Transformer-based models, the proposed method leverages the Shifted Window mechanism and hierarchical structure of Swin-Transformer to effectively detect synthetic images across multiple domains. It is trained and validated on the RVF10k and Real and Fake Face Detection datasets, demonstrating superior accuracy and generalization ability. By integrating image preprocessing and augmentation strategies, this study further enhances detection robustness, ensuring resilience to various forms of synthetic media. Experimental results confirm that the proposed model outperforms existing methods, particularly in detecting subtle manipulations in high-quality AI-generated images. Future research will extend its applicability to deepfake videos and synthetic text, further improving its generalization capabilities. By providing a scalable and efficient detection framework, this study contributes to the responsible and secure deployment of generative AI technologies.