Random perturbation subsampling for rank regression with massive data
摘要
Rank regression plays a fundamental role in statistical data analysis due to its robustness and high efficiency, and has been widely used in various scientific fields. However, when the data size is huge or massive, but the computing resource is limited, it can lead to an unacceptably computational cost for rank regression estimation. To handle this issue, this paper applies a random perturbation subsampling method to rank regression models. Specifically, we develop two repeatedly random perturbation subsampling algorithms to find the estimation of parameters. Two different weighting strategies with product weights and additive weights are examined in the objective function. Differing from the existing optimal and Poisson subsampling methods, our methods do not require the explicit calculation of subsampling probabilities for all data points, in which some unknown parameters often need to be estimated, thus making the implementation of our methods easier. Theoretically, statistical justifications are further provided for the proposed estimators including consistency and asymptotic normality. Extensive simulation studies and an empirical application are carried out to illustrate the effectiveness of the proposed methods.