Multimodal industrial anomaly detection has been extensively studied. However, when deployed in real-world industrial scenarios, it often encounters accuracy degradation due to modality missing, which stems from the high computational latency associated with processing certain modality, such as point clouds. To address the issue of modality missing, existing methods typically require a large number of labeled samples for model training. However, anomalous samples are scarce in industrial settings, making such approaches less practical. In this paper, we propose MissingClip, a novel framework for anomaly detection under modality missing based on multimodal large-scale model CLIP. MissingClip addresses modality missing by reconstructing semantic features that capture the missing information. First, we map 3D point clouds to 2D images, effectively supplementing the 2D modality with 3D information. Next, we extract both global and local features using a dual-attention mechanism. Second, we introduce a hybrid semantic prompt mechanism, which learns the normal and anomalous features of each modality through prompt texts, thereby recording modality features into the text modality. Finally, we propose a prompt reassembly mechanism that reconstructs missing modality features by combining the semantic features of prompts. Through extensive experiments and analysis, MissingClip demonstrates superior performance in industrial anomaly detection under missing modality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MissingClip: An Industrial Anomaly Detection Method Under Modality Missing

  • Tianyi Xu,
  • Ziqi Gan,
  • Xiaobo Zhou,
  • Fengbiao Zan,
  • Tie Qiu

摘要

Multimodal industrial anomaly detection has been extensively studied. However, when deployed in real-world industrial scenarios, it often encounters accuracy degradation due to modality missing, which stems from the high computational latency associated with processing certain modality, such as point clouds. To address the issue of modality missing, existing methods typically require a large number of labeled samples for model training. However, anomalous samples are scarce in industrial settings, making such approaches less practical. In this paper, we propose MissingClip, a novel framework for anomaly detection under modality missing based on multimodal large-scale model CLIP. MissingClip addresses modality missing by reconstructing semantic features that capture the missing information. First, we map 3D point clouds to 2D images, effectively supplementing the 2D modality with 3D information. Next, we extract both global and local features using a dual-attention mechanism. Second, we introduce a hybrid semantic prompt mechanism, which learns the normal and anomalous features of each modality through prompt texts, thereby recording modality features into the text modality. Finally, we propose a prompt reassembly mechanism that reconstructs missing modality features by combining the semantic features of prompts. Through extensive experiments and analysis, MissingClip demonstrates superior performance in industrial anomaly detection under missing modality.