<p>In this paper, we introduce a screening method for ultra-high-dimensional data with grouping structures. Our proposal is based on the improved projection correlation (IPC, for short), which effectively quantifies the dependence between two random vectors. The IPC-based group screening is model-free, robust to outliers or extreme values in the dataset, and computationally fast. We establish the ranking consistency and the sure screening properties for the IPC-based group screening procedure under mild assumptions. To specify the threshold of the screening procedure, we introduce group deep knockoffs, which improve the power of deep knockoffs in the context of group screening, and then advocate a two-step approach based on the group deep knockoffs. This approach ensures that the false discovery rate remains controlled below a predetermined level. Comprehensive simulations and an application to the Cancer Cell Line Encyclopedia (CCLE) RNAseq gene expression and transcript data demonstrate the finite sample performance of our approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-free group screening and FDR control with deep knockoffs

  • Jian Lin,
  • Liping Zhu,
  • Tingyou Zhou

摘要

In this paper, we introduce a screening method for ultra-high-dimensional data with grouping structures. Our proposal is based on the improved projection correlation (IPC, for short), which effectively quantifies the dependence between two random vectors. The IPC-based group screening is model-free, robust to outliers or extreme values in the dataset, and computationally fast. We establish the ranking consistency and the sure screening properties for the IPC-based group screening procedure under mild assumptions. To specify the threshold of the screening procedure, we introduce group deep knockoffs, which improve the power of deep knockoffs in the context of group screening, and then advocate a two-step approach based on the group deep knockoffs. This approach ensures that the false discovery rate remains controlled below a predetermined level. Comprehensive simulations and an application to the Cancer Cell Line Encyclopedia (CCLE) RNAseq gene expression and transcript data demonstrate the finite sample performance of our approach.