Energy-based causal disentanglement for compositional zero-shot learning
摘要
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen compositions of attributes and objects in images, leveraging prior knowledge of seen compositions of visual primitive concepts. Compositionality is considered a key capability for advancing artificial intelligence systems. Previous works in CZSL focus on embedding either the pair or the separate concepts using associative models. However, these methods often struggle to disentangle the entangled attributes and objects in images, failing to effectively capture the compositional relationships. To address this challenge, we propose a novel Energy-Based Causal Disentanglement (EBCD) model for compositional zero-shot learning. In the proposed EBCD model, energy-based models are used specifically for disentangling object and attribute representations in visual images. Additionally, causal interventions are employed to generate images with unseen compositions, ensuring that disentangled features remain stable under interventions. Extensive experiments are conducted on the public benchmark MIT-States and UT-Zappos datasets, and our method achieves better performance compared with previous state-of-the-art methods.