CLIPMisD: Few-Shot Prompt Learning for Misclassification Detection with Vision Language Model
摘要
Reliable prediction of classifier is essential for its deployment in high security situations. However, modern networks tend to be overconfident for misclassified predictions, urging confidence estimation to identify the erroneous of predictions. Despite the progress achieved by traditional neural networks on small-scale datasets, they all require training from scratch and there are no efficient and effective misclassification detection (MisD) methods, hindering practical application towards large-scale datasets. In this paper, we pave the way to exploit vision language model (VLM) leveraging text information to improve the performance of misclassification detection. Utilizing the power of VLM, we first construct prompt learning framework for MisD to refrain from training from scratch and therefore improve tuning efficiency. Then, we propose our approach named CLIPMisD, which uses textual guided negative augmentation along with a novel negative loss to mitigate the issue of overconfidence by pushing category prompts away from pseudo samples. We evaluate the performance of prompt learning methods and conduct comprehensive experiments on MisD task. Significant and consistent improvement demonstrates the effectiveness and superiority of our approach.