Clothing Purification with Causality Meets Vision-Language Pretraining Models
摘要
Vision-Language Pretraining (VLP) models have shown significant promise in scene understanding and representation learning. However, their application in fine-grained tasks like Cloth-Changing Person Re-Identification (CC-ReID) is challenging due to their reliance on unstable discriminative features such as clothing. Conversely, expert CC-ReID models possess exceptional fine-grained comprehension skills but struggle to obtain reliable cloth-agnostic representations, hindered by the time-consuming and labor-intensive process of obtaining precise annotations and the spurious data associations brought by the co-occurrence phenomenon of identity and clothing. This paper introduces the Causality-based Purification (CaPu) model. The CaPu constructs a clothing indication pipeline that leverages the unique strengths of multiple VLP models to efficiently capture clothing semantics. Utilizing these semantics, CaPu employs causality analysis to purify the relationship between learned visual features and intrinsic identity representation from two causal aspects: the Consistency Treatment Effect (CTE) and the Distinctiveness Treatment Effect (DTE). The CTE enhances feature consistency within each identity by simulating clothing changes. Meanwhile, the DTE enhances the model’s ability to perceive intrinsic identity representation. Extensive experiments on three standard CC-ReID datasets demonstrate that CaPu achieves state-of-the-art performance.