Face anti-spoofing (FAS) which serves to maintain the security and reliability of face recognition systems is a significant topic in face recognition. In recent years, Vision-Language models (e.g., CLIP) have been proved effective in the field of FAS. However, the current work has some limitations in terms of prompts. On the one hand, manually written prompts are often too simple, which limits the model’s in-depth mining of text supervision information and leads to poor generalization. On the other hand, although large language models can automatically generate richer prompts, their training and inference speed is slow, which may become a bottleneck in practical applications, affecting the usability and efficiency of the model. In this work, we propose a novel Cross-modal Feature Modulation (CMFM) method to adaptively generate image description from image itself and combines them with pre-designed text description to form the final prompt, which solves the inaccuracy of manual annotation. Meanwhile, we propose a Dynamic Adaptive Clustering (DAC) strategy designed to minimize intra-class distances among real samples while maximizing inter-class distances between real and spoof samples, thereby significantly enhancing the domain generalization of the model. The experimental results show our work significantly improves the generalization of models and outperforms the vast majority of state-of-the-art (SOTA) methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-Supervised Cross-Modal Feature Modulation and Dynamic Adaptive Clustering for Face Anti-Spoofing

  • Yulang Yuan

摘要

Face anti-spoofing (FAS) which serves to maintain the security and reliability of face recognition systems is a significant topic in face recognition. In recent years, Vision-Language models (e.g., CLIP) have been proved effective in the field of FAS. However, the current work has some limitations in terms of prompts. On the one hand, manually written prompts are often too simple, which limits the model’s in-depth mining of text supervision information and leads to poor generalization. On the other hand, although large language models can automatically generate richer prompts, their training and inference speed is slow, which may become a bottleneck in practical applications, affecting the usability and efficiency of the model. In this work, we propose a novel Cross-modal Feature Modulation (CMFM) method to adaptively generate image description from image itself and combines them with pre-designed text description to form the final prompt, which solves the inaccuracy of manual annotation. Meanwhile, we propose a Dynamic Adaptive Clustering (DAC) strategy designed to minimize intra-class distances among real samples while maximizing inter-class distances between real and spoof samples, thereby significantly enhancing the domain generalization of the model. The experimental results show our work significantly improves the generalization of models and outperforms the vast majority of state-of-the-art (SOTA) methods.