<p>This article proposes a novel fully Bayesian procedure to select the subset of regressors that are the relevant features explaining a target variable. We consider the normal linear regression model including the possible correlated design and the possible high-dimensional context where the number of explanatory variables might largely exceed the sample size. We generalize Zellner’s g-prior thanks to a random Wishart matrix and we present a straightforward stochastic search algorithm for posterior computation of the model parameters over all possible models generated. Particularly, we develop a Metropolis-Hastings-within-Gibbs scheme for the stochastic search to visit models having high posterior probabilities and to gather samples from the resulting posterior distributions. Using simulated and real datasets, we show that our methodology yields a higher frequency of selecting the correct variables and has a higher predictive power relative to other widely used variable selection methods such as adaptive Lasso, Bayesian adaptive Lasso and relative to well-known machine learning algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bayesian Adaptive Variable Selection with a Generalized g-prior

  • Djibril Ndiaye,
  • Khader Khadraoui

摘要

This article proposes a novel fully Bayesian procedure to select the subset of regressors that are the relevant features explaining a target variable. We consider the normal linear regression model including the possible correlated design and the possible high-dimensional context where the number of explanatory variables might largely exceed the sample size. We generalize Zellner’s g-prior thanks to a random Wishart matrix and we present a straightforward stochastic search algorithm for posterior computation of the model parameters over all possible models generated. Particularly, we develop a Metropolis-Hastings-within-Gibbs scheme for the stochastic search to visit models having high posterior probabilities and to gather samples from the resulting posterior distributions. Using simulated and real datasets, we show that our methodology yields a higher frequency of selecting the correct variables and has a higher predictive power relative to other widely used variable selection methods such as adaptive Lasso, Bayesian adaptive Lasso and relative to well-known machine learning algorithms.