<p>Face attack detection plays a vital role in safeguarding face recognition (<i>FR</i>) systems. Due to the unrestricted access to enormous face images and face manipulation tools circulating on the Internet, both physical and digital attacks pose significant threats to the widespread use of FR systems. However, previous works consider the detection of physical attacks and digital attacks as two independent tasks, causing an inferior generalization of attack detection among different categories. In this paper, we propose a unified multi-modal multi-scale hybrid CNN and transformer detection framework, namely <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\mathtt UniM2CT\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="monospace">U</mi> <mi>n</mi> <mi>i</mi> <mi>M</mi> <mn>2</mn> <mi>C</mi> <mi>T</mi> </mrow> </math></EquationSource> </InlineEquation>, for unified face attack detection. UniM2CT employs a dual-stream network architecture, comprising a spatial branch and a frequency domain branch. The spatial branch utilizes an EfficientNet to extract local CNN features, which are then fed into a proposed multi-scale ViT to capture global inconsistencies at multiple spatial levels. The frequency domain branch, on the other hand, employs a shallow CNN to learn the feature representations at low-, medium-, and high-frequency levels. A multi-modal fusion module is elaborately designed based on local, global, and self-attention mechanisms to integrate the three feature sources, enhancing the model’s ability to learn complementary face deception information. Experimental results demonstrate that <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\mathtt UniM2CT\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="monospace">U</mi> <mi>n</mi> <mi>i</mi> <mi>M</mi> <mn>2</mn> <mi>C</mi> <mi>T</mi> </mrow> </math></EquationSource> </InlineEquation> outperforms many state-of-the-art methods in unified face attack detection, exhibiting promising anti-attack capabilities. The source code is available at <a href="https://github.com/thirteenl/UniM2CT">https://github.com/thirteenl/UniM2CT</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unified face attack detection via multi-modal multi-scale CNN–ViT network: enhancing representation capability

  • Jun Liu,
  • Zengxi Huang,
  • Zehui Tang,
  • Tingsong Ma,
  • Shengke Zeng,
  • Yong Guo

摘要

Face attack detection plays a vital role in safeguarding face recognition (FR) systems. Due to the unrestricted access to enormous face images and face manipulation tools circulating on the Internet, both physical and digital attacks pose significant threats to the widespread use of FR systems. However, previous works consider the detection of physical attacks and digital attacks as two independent tasks, causing an inferior generalization of attack detection among different categories. In this paper, we propose a unified multi-modal multi-scale hybrid CNN and transformer detection framework, namely \(\mathtt UniM2CT\) U n i M 2 C T , for unified face attack detection. UniM2CT employs a dual-stream network architecture, comprising a spatial branch and a frequency domain branch. The spatial branch utilizes an EfficientNet to extract local CNN features, which are then fed into a proposed multi-scale ViT to capture global inconsistencies at multiple spatial levels. The frequency domain branch, on the other hand, employs a shallow CNN to learn the feature representations at low-, medium-, and high-frequency levels. A multi-modal fusion module is elaborately designed based on local, global, and self-attention mechanisms to integrate the three feature sources, enhancing the model’s ability to learn complementary face deception information. Experimental results demonstrate that \(\mathtt UniM2CT\) U n i M 2 C T outperforms many state-of-the-art methods in unified face attack detection, exhibiting promising anti-attack capabilities. The source code is available at https://github.com/thirteenl/UniM2CT.