<p>Face authentication methods are increasingly prone to sophisticated spoofing techniques, such as deepfakes and multi-modal forgeries, that are dangerous for applications with important security implications. Standard anti-spoofing techniques, though effective under pre-conditioned conditions, are not applicable to various real-world situations as they only take good-quality data and have weak temporal modeling capabilities. This paper presents a new multimodal approach using transformers for high quality temporal feature extraction and deformable convolutions for precise spatial analysis. The platform uses a fully integrated fusion method to incorporate RGB, depth, and binary masking to improve detection performance for different attacks and environments. Many experiments using benchmark datasets such as CelebA-Spoof, Replay-Attack and CASIA-FASD demonstrate the power of the framework, with an average accuracy of 99.1% and remarkably low HTER values. Cross-dataset testing validates its generalisability with very little performance loss on unseen data. The new approach substantially advances current practice by solving fundamental shortcomings of existing solutions including inflexibility in real-time environments and computational overhead. Its scale and real-time features make it ideal for use in IoT, banking and surveillance. This work provides a solid basis for safe and secure face authentication despite changing spoofing methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Transformer-Based Framework for Generalized Video Face Anti-Spoofing in Dynamic Environments

  • Devi Palanisamy,
  • S. Anila

摘要

Face authentication methods are increasingly prone to sophisticated spoofing techniques, such as deepfakes and multi-modal forgeries, that are dangerous for applications with important security implications. Standard anti-spoofing techniques, though effective under pre-conditioned conditions, are not applicable to various real-world situations as they only take good-quality data and have weak temporal modeling capabilities. This paper presents a new multimodal approach using transformers for high quality temporal feature extraction and deformable convolutions for precise spatial analysis. The platform uses a fully integrated fusion method to incorporate RGB, depth, and binary masking to improve detection performance for different attacks and environments. Many experiments using benchmark datasets such as CelebA-Spoof, Replay-Attack and CASIA-FASD demonstrate the power of the framework, with an average accuracy of 99.1% and remarkably low HTER values. Cross-dataset testing validates its generalisability with very little performance loss on unseen data. The new approach substantially advances current practice by solving fundamental shortcomings of existing solutions including inflexibility in real-time environments and computational overhead. Its scale and real-time features make it ideal for use in IoT, banking and surveillance. This work provides a solid basis for safe and secure face authentication despite changing spoofing methods.