In the era of Big Data, several design based subsampling methods are proposed to reduce costs (and time) and to help in informed decision making. Most of these approaches require the specification of a model. A wrong model assumption and/or the possible presence of outliers represent a limitation for the most commonly applied subsampling criteria. Through a simulation study, we explore if a subsampling method, originally introduced by [1] to avoid outliers, works well to account for model uncertainty and, on the other side, if the subsampling approach introduced by [2] to account for model misspecification, is robust to the presence of outliers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimal Subsampling from Big Datasets in Presence of Misspecification

  • Laura Deldossi,
  • Chiara Tommasi

摘要

In the era of Big Data, several design based subsampling methods are proposed to reduce costs (and time) and to help in informed decision making. Most of these approaches require the specification of a model. A wrong model assumption and/or the possible presence of outliers represent a limitation for the most commonly applied subsampling criteria. Through a simulation study, we explore if a subsampling method, originally introduced by [1] to avoid outliers, works well to account for model uncertainty and, on the other side, if the subsampling approach introduced by [2] to account for model misspecification, is robust to the presence of outliers.