<p>Large-scale foundation models, such as the contrastive language-image pre-training model and the align language model, have shown promising performance on downstream tasks. However, despite their accomplishments, these large-scale foundation models still exhibit limitations in handling certain out-of-distribution downstream tasks, especially in the field of few-shot domain adaptation (FSDA). Advanced works propose prompt learning to overcome the distribution shift. However, the existing methods mainly concentrate on learning <i>universal</i> prompts applicable across available domains, neglecting to learn <i>domain-specific</i> prompts for the target domain already known in FSDA tasks. To fill this gap, we propose a novel learning approach, termed <Emphasis Type="BoldItalic">en</Emphasis>tangle-<Emphasis Type="BoldItalic">t</Emphasis>hen-<Emphasis Type="BoldItalic">di</Emphasis>sentangle (EntDi), where each domain is assigned a distinct prompt to model the domain knowledge. The insight is that visual features from two domains, once entangled into a single representation, could be disentangled by leveraging domain-specific knowledge. Specifically, EntDi first entangles visual features from two images of disparate labels and domains. Subsequently, EntDi learns domain-specific prompts by predicting labels of these entangled features, where the labels are contingent on the domain-specific prompt used for prediction. Comprehensive experiments verify the efficacy of the proposed prompt learning approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Entangle-then-disentangle: a novel approach for enhancing large vision-language model

  • Jiajun Yuan,
  • Haiting Zheng,
  • Hang Yu,
  • Xiangfeng Luo

摘要

Large-scale foundation models, such as the contrastive language-image pre-training model and the align language model, have shown promising performance on downstream tasks. However, despite their accomplishments, these large-scale foundation models still exhibit limitations in handling certain out-of-distribution downstream tasks, especially in the field of few-shot domain adaptation (FSDA). Advanced works propose prompt learning to overcome the distribution shift. However, the existing methods mainly concentrate on learning universal prompts applicable across available domains, neglecting to learn domain-specific prompts for the target domain already known in FSDA tasks. To fill this gap, we propose a novel learning approach, termed entangle-then-disentangle (EntDi), where each domain is assigned a distinct prompt to model the domain knowledge. The insight is that visual features from two domains, once entangled into a single representation, could be disentangled by leveraging domain-specific knowledge. Specifically, EntDi first entangles visual features from two images of disparate labels and domains. Subsequently, EntDi learns domain-specific prompts by predicting labels of these entangled features, where the labels are contingent on the domain-specific prompt used for prediction. Comprehensive experiments verify the efficacy of the proposed prompt learning approach.