<p>In this paper, we propose to regularize non-symmetric correspondence analysis (NSCA) and its canonical variant by employing LASSO and group LASSO penalties. NSCA visualizes the asymmetric association structure of a categorical predictor variable and a categorical response variable through a biplot with points for the predictor categories and vectors for the response categories. In canonical NSCA, external information is available about the categories of the predictor variable and this information is used to linearly constrain the coordinates of the points. When the number of predictor categories is large or when the number of external variables is large, this leads to problems in terms of interpretation and/or estimation. To avoid these problems, we propose to use a LASSO or group LASSO penalty on the parameters. Such penalties shrink the parameters to zero, offering a sparse solution. Therefore, we first cast (constrained) NSCA as a least squares estimation problem and then add the penalty to the least squares loss function. We derive a Majorization-Minimization algorithm to minimize this loss function. A bootstrap procedure is proposed for model selection, that is, determining the optimal dimensionality and optimal value of the penalty parameter. The procedures are illustrated using two empirical data sets, one for constrained (i.e., canonical) NSCA, and one for unconstrained NSCA. We discuss in detail the model selection procedure and the interpretation of the selected model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sparse constrained and unconstrained non-symmetric correspondence analysis

  • Mark de Rooij,
  • Rosaria Lombardo

摘要

In this paper, we propose to regularize non-symmetric correspondence analysis (NSCA) and its canonical variant by employing LASSO and group LASSO penalties. NSCA visualizes the asymmetric association structure of a categorical predictor variable and a categorical response variable through a biplot with points for the predictor categories and vectors for the response categories. In canonical NSCA, external information is available about the categories of the predictor variable and this information is used to linearly constrain the coordinates of the points. When the number of predictor categories is large or when the number of external variables is large, this leads to problems in terms of interpretation and/or estimation. To avoid these problems, we propose to use a LASSO or group LASSO penalty on the parameters. Such penalties shrink the parameters to zero, offering a sparse solution. Therefore, we first cast (constrained) NSCA as a least squares estimation problem and then add the penalty to the least squares loss function. We derive a Majorization-Minimization algorithm to minimize this loss function. A bootstrap procedure is proposed for model selection, that is, determining the optimal dimensionality and optimal value of the penalty parameter. The procedures are illustrated using two empirical data sets, one for constrained (i.e., canonical) NSCA, and one for unconstrained NSCA. We discuss in detail the model selection procedure and the interpretation of the selected model.