ICR-Net: Robust Deepfake Detection Under Temporal Corruption
摘要
Deepfake video detection aims to distinguish AI-generated facial forgeries from authentic videos. While recent methods show strong performance under spatial corruptions, their robustness against temporal corruptions remains largely unexplored. In practical streaming scenarios, network disruptions such as packet loss, bit errors, and aggressive compression induce temporal degradations that current benchmarks do not cover. To bridge this gap, we introduce the DeepFake Temporal Corruption Benchmark (DF-TCB), built on FaceForensics++ and DFDC datasets with diverse corruption types and severities. Our analysis reveals that existing detectors are highly fragile under these disruptions. We therefore propose ICR-Net, which predicts frame reliability and selectively corrects corrupted features. By leveraging clean-corrupted contrastive learning, it extracts corruption-invariant, class-separable representations. We demonstrate that ICR-Net achieves state-of-the-art robustness and cross-dataset generalization under diverse temporal corruptions.