Face mask detection has become crucial for public health and safety, especially during the COVID-19 pandemic. The existing methods, relying on large datasets of labeled human faces, pose privacy concerns and may not achieve high accuracy in diverse environments. In this paper, we present an innovative approach namely VICL-CLIP, which incorporates the Visual In-Context Learning (V-ICL) paradigm into the CLIP model to enhance face mask detection. By leveraging standardized cartoon images as learning context, our method addresses privacy issues while it also significantly improves detection accuracy. Specifically, we design effective multimodal prompts for in-context learning. Cartoon images with and without masks are proposed as the image prompts, while their corresponding text prompts are curated as the positive and negative contexts for the CLIP model. In this way, the model is able to be refined to generalize the capability from abstract representations to real human faces, through the inherent visual-text linkage. Our extensive experiments were conducted based on an real-world COVID Face Mask Detection Dataset. Our VICL-CLIP model achieves an excellent detection accuracy of 97%, outperforming all conventional methods and other state-of-the-art models. Moreover, this work underscores the potential of integrating the V-ICL learning paradigm into powerful vision-language foundation models to improve the mask detection accuracy while preserving privacy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VICL-CLIP: Enhancing Face Mask Detection in Context with Multimodal Foundation Models

  • Xinyi Gao,
  • Yanbin Liu,
  • Minh Nguyen,
  • Wei Qi Yan

摘要

Face mask detection has become crucial for public health and safety, especially during the COVID-19 pandemic. The existing methods, relying on large datasets of labeled human faces, pose privacy concerns and may not achieve high accuracy in diverse environments. In this paper, we present an innovative approach namely VICL-CLIP, which incorporates the Visual In-Context Learning (V-ICL) paradigm into the CLIP model to enhance face mask detection. By leveraging standardized cartoon images as learning context, our method addresses privacy issues while it also significantly improves detection accuracy. Specifically, we design effective multimodal prompts for in-context learning. Cartoon images with and without masks are proposed as the image prompts, while their corresponding text prompts are curated as the positive and negative contexts for the CLIP model. In this way, the model is able to be refined to generalize the capability from abstract representations to real human faces, through the inherent visual-text linkage. Our extensive experiments were conducted based on an real-world COVID Face Mask Detection Dataset. Our VICL-CLIP model achieves an excellent detection accuracy of 97%, outperforming all conventional methods and other state-of-the-art models. Moreover, this work underscores the potential of integrating the V-ICL learning paradigm into powerful vision-language foundation models to improve the mask detection accuracy while preserving privacy.