Entropy-Based Subsampling Methods for Big Data
摘要
Under big data settings, parameter estimation can often become computationally infeasible for popular regression models. Subsampling techniques, in such cases, play a significant role in better computational efficiency with only a slight loss of statistical estimation accuracy. For multiple linear regression models, traditional subsampling techniques like leveraging methods have primarily measured information loss based only on covariates (design matrix) excluding the responses and thus often lead to a loss of statistical estimation accuracy. Unlike the idea of only keeping the subsample, the problem is viewed as extracting a representative set of the full data in terms of entropy. Two naive methods are proposed, both inspired by the benchmark algorithm information-based optimal subsample selection (IBOSS). One method is based on an assumed (known) subsample size, which we refer to as likelihood-based optimal subsample selection (LBOSS). The other method automatically determines the subsample size and is named Bayesian-based optimal subsample selection (BBOSS). The proposed entropy-based criteria not only provide a better measure of the information loss due to subsampling but are also applicable to any likelihood-based estimation methods. In addition to some theoretical guarantees of the proposed methods, we provide extensive numerical illustrations to compare the proposed methods with some of the recently published methods in the literature. In terms of computational efficiency and statistical accuracy, the proposed methods are shown to perform relatively better.