As deepfake technology advances rapidly, face forgery has become increasingly complex, posing significant security and privacy risks. Current CLIP-based detectors primarily extract general visual features, which are not specifically designed to identify forgery artifacts. This limitation prevents the detectors from effectively learning forgery-specific features. To address these limitations, we propose a novel forgery-aware adaptive CLIP with two key contributions: a Hierarchical Visual Feature Adapter (HVFA) and a Latent Prompt Learning (LPL). The HVFA customizes CLIP by refining its hierarchical visual features to detect forgery clues at different levels. Meanwhile, the LPL enhances the model’s adaptability by incorporating a learnable, forgery-specific embedding with predefined prompt embeddings in the latent space. By optimizing both the vision and language branches of CLIP, our method effectively adapts the model for detecting forged faces. Extensive experiments demonstrate that our method outperforms current state-of-the-art methods in terms of generalization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Forgery-Aware Adaptive CLIP for Generalizable Face Forgery Detection

  • Xiaoning Li,
  • Xinyu Chen,
  • Yongcun Zhang,
  • Yingxin Lai,
  • Guimin Shi,
  • Zhiming Luo

摘要

As deepfake technology advances rapidly, face forgery has become increasingly complex, posing significant security and privacy risks. Current CLIP-based detectors primarily extract general visual features, which are not specifically designed to identify forgery artifacts. This limitation prevents the detectors from effectively learning forgery-specific features. To address these limitations, we propose a novel forgery-aware adaptive CLIP with two key contributions: a Hierarchical Visual Feature Adapter (HVFA) and a Latent Prompt Learning (LPL). The HVFA customizes CLIP by refining its hierarchical visual features to detect forgery clues at different levels. Meanwhile, the LPL enhances the model’s adaptability by incorporating a learnable, forgery-specific embedding with predefined prompt embeddings in the latent space. By optimizing both the vision and language branches of CLIP, our method effectively adapts the model for detecting forged faces. Extensive experiments demonstrate that our method outperforms current state-of-the-art methods in terms of generalization.