Abstract <p>This paper deals with the problem of nonparametric regression when the response variable may be missing but not necessarily at random. Here, we propose a new approach to construct kernel-type estimators of an unknown regression function based on Horvitz–Thompson inverse weighting when the data suffers from missing response values. The proposed approach may be viewed as a two-step procedure: the first step involves constructing a family of kernel-type regression estimators based on inverse weighting where the members of this family are indexed by the unknown parameters of the missing probability mechanism (the selection probability). In the second step, a search will be carried out to find the member of a cover of this family that has the smallest mean-squared prediction error. Furthermore, we establish exponential performance bounds on the deviations of the proposed estimators from the true regression curve in general <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(L_{p}\)</EquationSource> <!--MMStat2570013Mojirsheibani-m3--> </InlineEquation> norms; these bounds yield various strong convergence results. We also study the rates of convergence of these estimators. As an important application of our results, we consider the problem of statistical classification with incomplete data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On Nonparametric Regression with Incomplete Data via Inverse-Weighting and Their Convergence in \(\boldsymbol{L_{p}}\) Norms

  • Majid Mojirsheibani,
  • Tatiana Shahinian

摘要

Abstract

This paper deals with the problem of nonparametric regression when the response variable may be missing but not necessarily at random. Here, we propose a new approach to construct kernel-type estimators of an unknown regression function based on Horvitz–Thompson inverse weighting when the data suffers from missing response values. The proposed approach may be viewed as a two-step procedure: the first step involves constructing a family of kernel-type regression estimators based on inverse weighting where the members of this family are indexed by the unknown parameters of the missing probability mechanism (the selection probability). In the second step, a search will be carried out to find the member of a cover of this family that has the smallest mean-squared prediction error. Furthermore, we establish exponential performance bounds on the deviations of the proposed estimators from the true regression curve in general \(L_{p}\) norms; these bounds yield various strong convergence results. We also study the rates of convergence of these estimators. As an important application of our results, we consider the problem of statistical classification with incomplete data.