<p>Unsupervised domain adaptation (UDA) transfers knowledge from a labeled source domain to an unlabeled target domain by jointly optimizing source classification and domain-alignment objectives. A key practical challenge is selecting the trade-off coefficient <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\lambda \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>λ</mi> </math></EquationSource> </InlineEquation>, which controls the strength of domain alignment. Fixed values and manually designed schedules are commonly used, but the appropriate balance can vary across datasets, transfer directions, and training stages. This paper proposes Adaptive Gradient-Norm Weighting (AWD), a lightweight scalarization mechanism that dynamically adjusts <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\lambda \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>λ</mi> </math></EquationSource> </InlineEquation> using the ratio between the gradient norms of the classification and alignment losses. AWD introduces no additional learnable parameters and can be applied to UDA objectives of the form <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\mathcal {L}_{\textrm{CE}}+\lambda \mathcal {L}_{\textrm{Align}}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi mathvariant="script">L</mi> <mtext>CE</mtext> </msub> <mo>+</mo> <mi>λ</mi> <msub> <mi mathvariant="script">L</mi> <mtext>Align</mtext> </msub> </mrow> </math></EquationSource> </InlineEquation>. We evaluate AWD within Deep Adaptation Network (DAN) and Domain-Adversarial Neural Network (DANN) on Office-31, Office-Home, DomainNet, and digit adaptation benchmarks using five independent seeds. Across the evaluated benchmarks, AWD consistently improves average performance over fixed and scheduled baselines, with the strongest gains on more difficult transfer settings. On Office-Home with DANN, AWD improves average accuracy over fixed <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\lambda {=}1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>λ</mi> <mo>=</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation> by <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(+4.4\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>+</mo> <mn>4.4</mn> </mrow> </math></EquationSource> </InlineEquation> percentage points, while avoiding manual per-task grid search. The method incurs only modest computational overhead. Overall, the results show that gradient-aware adaptive weighting is a simple, practical, and interpretable mechanism for balancing classification and alignment losses in UDA.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive gradient-norm weighting for improved domain adversarial training

  • Iman Khazrak,
  • Mohammadhossein Homaei,
  • Mostafa M. Rezaee,
  • Robert C. Green II

摘要

Unsupervised domain adaptation (UDA) transfers knowledge from a labeled source domain to an unlabeled target domain by jointly optimizing source classification and domain-alignment objectives. A key practical challenge is selecting the trade-off coefficient \(\lambda \) λ , which controls the strength of domain alignment. Fixed values and manually designed schedules are commonly used, but the appropriate balance can vary across datasets, transfer directions, and training stages. This paper proposes Adaptive Gradient-Norm Weighting (AWD), a lightweight scalarization mechanism that dynamically adjusts \(\lambda \) λ using the ratio between the gradient norms of the classification and alignment losses. AWD introduces no additional learnable parameters and can be applied to UDA objectives of the form \(\mathcal {L}_{\textrm{CE}}+\lambda \mathcal {L}_{\textrm{Align}}\) L CE + λ L Align . We evaluate AWD within Deep Adaptation Network (DAN) and Domain-Adversarial Neural Network (DANN) on Office-31, Office-Home, DomainNet, and digit adaptation benchmarks using five independent seeds. Across the evaluated benchmarks, AWD consistently improves average performance over fixed and scheduled baselines, with the strongest gains on more difficult transfer settings. On Office-Home with DANN, AWD improves average accuracy over fixed \(\lambda {=}1\) λ = 1 by \(+4.4\) + 4.4 percentage points, while avoiding manual per-task grid search. The method incurs only modest computational overhead. Overall, the results show that gradient-aware adaptive weighting is a simple, practical, and interpretable mechanism for balancing classification and alignment losses in UDA.