Vision Transformer (ViT) contributes to accurate change detection with robustness to background changes. However, retraining ViT requires a large amount of computation to adapt to unlearned scenes. This study investigates the addition of learnable parameters into change detection ViT to reduce the computational complexity of retraining. We introduce MLP as an adapter as an addition to the attention output and the residual connection of the change detection ViT and apply LoRA method to the change detection ViT. We evaluate the retraining of additional parameter models for various background changes and analyze proper setting of additional parameters to adapt the target scenes. Introducing MLP and LoRA to change detection ViT improves the accuracy for the target scenes without competition between two additional parameter methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of Adapter in Attention of Change Detection Vision Transformer

  • Ryunosuke Hamada,
  • Tsubasa Minematsu,
  • Cheng Tang,
  • Atsushi Shimada

摘要

Vision Transformer (ViT) contributes to accurate change detection with robustness to background changes. However, retraining ViT requires a large amount of computation to adapt to unlearned scenes. This study investigates the addition of learnable parameters into change detection ViT to reduce the computational complexity of retraining. We introduce MLP as an adapter as an addition to the attention output and the residual connection of the change detection ViT and apply LoRA method to the change detection ViT. We evaluate the retraining of additional parameter models for various background changes and analyze proper setting of additional parameters to adapt the target scenes. Introducing MLP and LoRA to change detection ViT improves the accuracy for the target scenes without competition between two additional parameter methods.