Optimal Subsampling from Big Datasets in Presence of Misspecification
摘要
In the era of Big Data, several design based subsampling methods are proposed to reduce costs (and time) and to help in informed decision making. Most of these approaches require the specification of a model. A wrong model assumption and/or the possible presence of outliers represent a limitation for the most commonly applied subsampling criteria. Through a simulation study, we explore if a subsampling method, originally introduced by [1] to avoid outliers, works well to account for model uncertainty and, on the other side, if the subsampling approach introduced by [2] to account for model misspecification, is robust to the presence of outliers.