Distribution shifts during test-time are prevalent in most machine learning applications and can often lead to a significant decline in the model’s performance. Foundation models like CLIP demonstrate zero-shot capabilities across various distributions. Additional fine-tuning on a specific dataset increases task performance but often reduces robustness to distribution shifts. Linearly interpolating the weights of the zero-shot and fine-tuned models (WiSE-FT) improves generalization capabilities, while maintaining task performance. Paradigms like online test-time adaptation (TTA) and test-time training (TTT) address distributional shifts by continuously updating the model during test-time. In light of these findings, we propose a novel method—adaptive weight-space ensembling (AdaWiSE) that dynamically balances between task specialization and generalization by adaptively interpolating the zero-shot and fine-tuned model during test-time. AdaWiSE utilizes Bayesian optimization to effectively determine the currently optimal mixing coefficient for interpolation, based on minimizing the entropy of the model’s predictions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Fine-Tuning and Adaptation of Zero-Shot Models via Adaptive Weight-Space Ensembling

  • Mario Döbler,
  • Michael Feil,
  • Robert A. Marsden,
  • Bin Yang

摘要

Distribution shifts during test-time are prevalent in most machine learning applications and can often lead to a significant decline in the model’s performance. Foundation models like CLIP demonstrate zero-shot capabilities across various distributions. Additional fine-tuning on a specific dataset increases task performance but often reduces robustness to distribution shifts. Linearly interpolating the weights of the zero-shot and fine-tuned models (WiSE-FT) improves generalization capabilities, while maintaining task performance. Paradigms like online test-time adaptation (TTA) and test-time training (TTT) address distributional shifts by continuously updating the model during test-time. In light of these findings, we propose a novel method—adaptive weight-space ensembling (AdaWiSE) that dynamically balances between task specialization and generalization by adaptively interpolating the zero-shot and fine-tuned model during test-time. AdaWiSE utilizes Bayesian optimization to effectively determine the currently optimal mixing coefficient for interpolation, based on minimizing the entropy of the model’s predictions.