VICL-CLIP: Enhancing Face Mask Detection in Context with Multimodal Foundation Models
摘要
Face mask detection has become crucial for public health and safety, especially during the COVID-19 pandemic. The existing methods, relying on large datasets of labeled human faces, pose privacy concerns and may not achieve high accuracy in diverse environments. In this paper, we present an innovative approach namely VICL-CLIP, which incorporates the Visual In-Context Learning (V-ICL) paradigm into the CLIP model to enhance face mask detection. By leveraging standardized cartoon images as learning context, our method addresses privacy issues while it also significantly improves detection accuracy. Specifically, we design effective multimodal prompts for in-context learning. Cartoon images with and without masks are proposed as the image prompts, while their corresponding text prompts are curated as the positive and negative contexts for the CLIP model. In this way, the model is able to be refined to generalize the capability from abstract representations to real human faces, through the inherent visual-text linkage. Our extensive experiments were conducted based on an real-world COVID Face Mask Detection Dataset. Our VICL-CLIP model achieves an excellent detection accuracy of 97%, outperforming all conventional methods and other state-of-the-art models. Moreover, this work underscores the potential of integrating the V-ICL learning paradigm into powerful vision-language foundation models to improve the mask detection accuracy while preserving privacy.