<p>Model selection criteria constitute one of the most critical issues in data clustering via mixture modeling. In this framework, criteria based on penalized log-likelihood (e.g. AIC and BIC) are usually adopted and one single model is finally selected. In this paper, an approach taking into account uncertainty in model selection is proposed in the maximum likelihood framework yielding a conditional probability distribution on the number of components, given a sample. The assessment of the number of population components, which are referred to as underlying populations, is also proposed based on employing a non-parametric bootstrap sampling of the observed data sample. Finally, a novel clustering approach relying on the developed concepts is presented. The proposal is illustrated on the ground of a numerical study based on both simulated and real data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncertain model selection criteria for mixture modeling

  • Volodymyr Melnykov,
  • Salvatore Ingrassia

摘要

Model selection criteria constitute one of the most critical issues in data clustering via mixture modeling. In this framework, criteria based on penalized log-likelihood (e.g. AIC and BIC) are usually adopted and one single model is finally selected. In this paper, an approach taking into account uncertainty in model selection is proposed in the maximum likelihood framework yielding a conditional probability distribution on the number of components, given a sample. The assessment of the number of population components, which are referred to as underlying populations, is also proposed based on employing a non-parametric bootstrap sampling of the observed data sample. Finally, a novel clustering approach relying on the developed concepts is presented. The proposal is illustrated on the ground of a numerical study based on both simulated and real data.