<p>A notable issue in the semi-supervised logistic regression estimation stems from the scarcity of available labeled data. To solve this problem, we propose a novel approach to enhance the estimation of high-dimensional logistic regression by leveraging network information. By formulating the total effect of covariates on network connections and establishing the relationship between class labels and network connections, we deduce that the regression coefficients fall into a low-dimensional subspace under mild conditions, transforming the high-dimensional estimation problem into a low-dimensional one. Particularly, we construct efficient estimators of parameters without the sparsity assumption, allowing the dimension to be larger than the sample size of labeled data. Theoretically, it is shown that the proposed estimators enjoy a faster convergence rate than existing methods using labeled data only, for both dense and sparse networks. Additionally, we extend the idea to the network with community structure, in which both the parameter estimation and community detection are simultaneously explored. Extensive simulation studies and real data analysis confirm the advantages of the proposed method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Network-assisted Semi-supervised Logistic Regression

  • Shengbin Zheng,
  • Zhonghan Wang,
  • Junlong Zhao

摘要

A notable issue in the semi-supervised logistic regression estimation stems from the scarcity of available labeled data. To solve this problem, we propose a novel approach to enhance the estimation of high-dimensional logistic regression by leveraging network information. By formulating the total effect of covariates on network connections and establishing the relationship between class labels and network connections, we deduce that the regression coefficients fall into a low-dimensional subspace under mild conditions, transforming the high-dimensional estimation problem into a low-dimensional one. Particularly, we construct efficient estimators of parameters without the sparsity assumption, allowing the dimension to be larger than the sample size of labeled data. Theoretically, it is shown that the proposed estimators enjoy a faster convergence rate than existing methods using labeled data only, for both dense and sparse networks. Additionally, we extend the idea to the network with community structure, in which both the parameter estimation and community detection are simultaneously explored. Extensive simulation studies and real data analysis confirm the advantages of the proposed method.