<p>With the steady advancement of data collection technologies, practical users of statistical methods increasingly turn to the mixture of factor analysers (MFA) for model-based clustering and dimensionality reduction. However, as the number of measurements grows, so does the likelihood of missing data and outliers, which can lead to biased parameter estimates, reduced stability and robustness, and ultimately inaccurate inferences. This paper presents a new variant of MFA model that can accommodate missing data and mild outliers. The main assumption of the proposed model is that the latent factors and idiosyncratic errors follow jointly a contaminated-normal distribution, which incorporates parameters for automatic outlier detection. We develop the ECM and AECM algorithms to compute maximum likelihood parameter estimates. Asymptotic standard errors of parameters are derived by offering an information-based approach. Several simulation experiments are conducted to examine the asymptotic properties of the ML estimators and assess the model’s ability to mitigate the influence of missing data and outliers. We further illustrate the model’s practical applicability in social data analysis and image reconstruction, using cost-of-living data and the Barbara image as case studies. Software implementing the presented methodology is available at <a href="https://github.com/leila-shahriari/CNMFA-Model">https://github.com/leila-shahriari/CNMFA-Model</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-based clustering of high-dimensional incomplete data via contaminated-normal mixtures

  • Leila Shahriari,
  • Mehrdad Naderi,
  • Mohsen Khosravi

摘要

With the steady advancement of data collection technologies, practical users of statistical methods increasingly turn to the mixture of factor analysers (MFA) for model-based clustering and dimensionality reduction. However, as the number of measurements grows, so does the likelihood of missing data and outliers, which can lead to biased parameter estimates, reduced stability and robustness, and ultimately inaccurate inferences. This paper presents a new variant of MFA model that can accommodate missing data and mild outliers. The main assumption of the proposed model is that the latent factors and idiosyncratic errors follow jointly a contaminated-normal distribution, which incorporates parameters for automatic outlier detection. We develop the ECM and AECM algorithms to compute maximum likelihood parameter estimates. Asymptotic standard errors of parameters are derived by offering an information-based approach. Several simulation experiments are conducted to examine the asymptotic properties of the ML estimators and assess the model’s ability to mitigate the influence of missing data and outliers. We further illustrate the model’s practical applicability in social data analysis and image reconstruction, using cost-of-living data and the Barbara image as case studies. Software implementing the presented methodology is available at https://github.com/leila-shahriari/CNMFA-Model.