Spatial-Temporal Graph Learning for DeepFake Video Detection
摘要
Face forgery by DeepFake is widespread on the Internet and has raised severe social concerns. This paper proposes a novel approach called spatial-temporal graph learning (STGL). This framework maximizes the potential of graph convolutional network (GCN) by leveraging the principle of message passing and focuses on relationship learning through purposefully designed graph structures. Specifically, we devised both local and global spatial-temporal graph structures, where nodes represent patches and frames separately and the relations stand as edges. These graph structures are the basis of relation learning in local and global branches. Finally, the output features of the two branches are combined to obtain more comprehensive and meaningful spatial-temporal features. Through rigorous experimentation across diverse datasets, we demonstrate that the proposed method surpasses state-of-the-art face forgery detection techniques and maintains efficacy when applied to unseen forgery datasets.