Entangle-then-disentangle: a novel approach for enhancing large vision-language model
摘要
Large-scale foundation models, such as the contrastive language-image pre-training model and the align language model, have shown promising performance on downstream tasks. However, despite their accomplishments, these large-scale foundation models still exhibit limitations in handling certain out-of-distribution downstream tasks, especially in the field of few-shot domain adaptation (FSDA). Advanced works propose prompt learning to overcome the distribution shift. However, the existing methods mainly concentrate on learning universal prompts applicable across available domains, neglecting to learn domain-specific prompts for the target domain already known in FSDA tasks. To fill this gap, we propose a novel learning approach, termed