<p>High-dimensional variable selection under domain adaptation poses a critical challenge in modern machine learning. Traditional selection methods often fail to generalize when the training (source) and testing (target) distributions differ, particularly in scenarios where the target domain lacks labeled data. In this paper, we propose a novel method that integrates transferable representation learning, model-based sensitivity analysis, and statistical stability selection to identify variables that are both predictive and transferable. Specifically, we design a three-stage framework that (i) learns domain-invariant embeddings using distribution alignment, (ii) computes variable importance through gradient attribution and subsample-based stability analysis, and (iii) fuses these perspectives into a robust variable selection strategy. We further formulate the entire process within a rigorous mathematical framework and provide theoretical guarantees, including a generalization bound for target risk, stability-based false selection control, and information retention of selected variables. Experimental results on both synthetic and real-world datasets demonstrate that our approach outperforms existing methods in terms of both predictive accuracy and variable interpretability across domains.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domain-adaptive high-dimensional variable selection via transfer learning and stability fusion

  • Yihao Wang

摘要

High-dimensional variable selection under domain adaptation poses a critical challenge in modern machine learning. Traditional selection methods often fail to generalize when the training (source) and testing (target) distributions differ, particularly in scenarios where the target domain lacks labeled data. In this paper, we propose a novel method that integrates transferable representation learning, model-based sensitivity analysis, and statistical stability selection to identify variables that are both predictive and transferable. Specifically, we design a three-stage framework that (i) learns domain-invariant embeddings using distribution alignment, (ii) computes variable importance through gradient attribution and subsample-based stability analysis, and (iii) fuses these perspectives into a robust variable selection strategy. We further formulate the entire process within a rigorous mathematical framework and provide theoretical guarantees, including a generalization bound for target risk, stability-based false selection control, and information retention of selected variables. Experimental results on both synthetic and real-world datasets demonstrate that our approach outperforms existing methods in terms of both predictive accuracy and variable interpretability across domains.