Multi-target Contrastive Objective for Learning Property-Aware Vision-Language Representation
摘要
Learning aligned image-text representations in a shared embedding space is crucial for vision-language tasks. While contrastive loss is highly effective for representation learning, it typically binds one image to one text caption, limiting the ability to capture individual properties from the text. Diagnostic datasets and benchmarks are vital for developing and understanding multi-property vision-language representations but are currently underexplored. We introduce Prop-Clevr, a diagnostic benchmark designed to evaluate the ability of models to capture multiple properties in vision-language representations. We propose a novel training objective, multi-target contrastive loss, which aligns image embeddings with multiple properties in the corresponding text. Empirical experiments on Prop-Clevr in image classification and text-image retrieval tasks demonstrate the effectiveness of our training objective in producing high-quality embeddings that capture individual properties in the images and texts. We provide the assets, codes, and tools that allow for high customization of Prop-Clevr, enabling the creation of new benchmarks to study and diagnose vision-language and related tasks.