<p>The hidden confounded model has been widely applied in many fields. Without adjusting for the hidden confounders, the estimators from the standard methods of high-dimensional models could be biased, potentially leading to spurious scientific discoveries. Meanwhile, distributed computation has attracted wide attention in modern statistical learning. Based on high-dimensional confounded models, this paper proposes a deconfounding and debiasing approach for distributed computing, aiming to obtain accurate estimation by reducing the confounding effect and bias. Two different distributed methods are applied: one is the straightforward divide-and-conquer (DC) method and the other is the communication-efficient surrogate likelihood (CSL) method. The former is easy to use in practice, while the latter uses surrogate loss to achieve better performance than the former through multiple iterations. The estimation accuracy and asymptotic theories for both DC and CSL estimators are established. Extensive simulation experiments verify the good performance of the two estimators and two real data applications are also presented to illustrate their validity and feasibility.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distributed Estimation and Inference for High-Dimensional Confounded Models

  • Jin Liu,
  • Yuxin Fei,
  • Wei Ma,
  • Lei Wang

摘要

The hidden confounded model has been widely applied in many fields. Without adjusting for the hidden confounders, the estimators from the standard methods of high-dimensional models could be biased, potentially leading to spurious scientific discoveries. Meanwhile, distributed computation has attracted wide attention in modern statistical learning. Based on high-dimensional confounded models, this paper proposes a deconfounding and debiasing approach for distributed computing, aiming to obtain accurate estimation by reducing the confounding effect and bias. Two different distributed methods are applied: one is the straightforward divide-and-conquer (DC) method and the other is the communication-efficient surrogate likelihood (CSL) method. The former is easy to use in practice, while the latter uses surrogate loss to achieve better performance than the former through multiple iterations. The estimation accuracy and asymptotic theories for both DC and CSL estimators are established. Extensive simulation experiments verify the good performance of the two estimators and two real data applications are also presented to illustrate their validity and feasibility.