Domain adaptation (DA) aims to transfer knowledge from labeled source domains to unlabeled target domains, addressing the challenge of model generalization when there is a distribution mismatch between training and testing data. While many Vision Transformer (ViT)-based methods have been developed for DA, they focus primarily on improving accuracy, with less emphasis on accelerating inference on unlabeled target domains. In this paper, we propose a novel method named Cascaded Adaptive Vision Transformer (CAViT), which dynamically adjusts token counts for each input image by cascading multiple transformers with increasing tokens. During testing, “easier” images exit early, while “harder” images are processed further until confident predictions are achieved. We further enhance domain adversarial learning by incorporating a token-level domain discriminator in the attention layer, which assigns distinct weights to different patch tokens. This enables the network to learn features with cross-domain transferability and discriminative capabilities, achieving effective feature alignment. Experimental results demonstrate that our method not only improves accuracy but also significantly reduces computational costs, as evidenced by results on three benchmark datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerating Domain Adaptation with Cascaded Adaptive Vision Transformer

  • Qilin Jiang,
  • Chaoran Cui,
  • Chunyun Zhang,
  • Yongrui Zhen,
  • Shuai Gong,
  • Ziyi Liu,
  • Fan’an Meng,
  • Hongyan Zhao

摘要

Domain adaptation (DA) aims to transfer knowledge from labeled source domains to unlabeled target domains, addressing the challenge of model generalization when there is a distribution mismatch between training and testing data. While many Vision Transformer (ViT)-based methods have been developed for DA, they focus primarily on improving accuracy, with less emphasis on accelerating inference on unlabeled target domains. In this paper, we propose a novel method named Cascaded Adaptive Vision Transformer (CAViT), which dynamically adjusts token counts for each input image by cascading multiple transformers with increasing tokens. During testing, “easier” images exit early, while “harder” images are processed further until confident predictions are achieved. We further enhance domain adversarial learning by incorporating a token-level domain discriminator in the attention layer, which assigns distinct weights to different patch tokens. This enables the network to learn features with cross-domain transferability and discriminative capabilities, achieving effective feature alignment. Experimental results demonstrate that our method not only improves accuracy but also significantly reduces computational costs, as evidenced by results on three benchmark datasets.