<p>Few clustering methods suit ordinal data, leading many researchers to incorrectly treat ordinal variables as continuous or nominal, which weakens statistical analysis and inference. Likelihood-based methods like finite mixtures can be used with ordinal-specific models, enabling statistical inference and model selection. The novelty of this article lies in extending the model-based clustering structure for the adjacent-categories logit model, which has been commonly used in ordinal regression but not, up to now, for clustering. Our data matrix has subjects as rows, and a set of ordinal responses, such as survey question responses, as columns. We cluster the subjects (rows) and/or questions (columns) via finite mixtures, using the expectation–maximization (EM) algorithm to estimate parameters. We assess the performance of the parameter estimates and test the effectiveness of the information-based asymptotic approximation to the standard errors of the parameter estimators via simulations. Additionally, we illustrate our example with a real-world linguistics dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Model-based Clustering Using Adjacent-Categories Logit Models via Finite-Mixtures for Ordinal Data

  • Ying Cui,
  • Lingyu Li,
  • Louise McMillan,
  • Richard Arnold,
  • Ivy Liu

摘要

Few clustering methods suit ordinal data, leading many researchers to incorrectly treat ordinal variables as continuous or nominal, which weakens statistical analysis and inference. Likelihood-based methods like finite mixtures can be used with ordinal-specific models, enabling statistical inference and model selection. The novelty of this article lies in extending the model-based clustering structure for the adjacent-categories logit model, which has been commonly used in ordinal regression but not, up to now, for clustering. Our data matrix has subjects as rows, and a set of ordinal responses, such as survey question responses, as columns. We cluster the subjects (rows) and/or questions (columns) via finite mixtures, using the expectation–maximization (EM) algorithm to estimate parameters. We assess the performance of the parameter estimates and test the effectiveness of the information-based asymptotic approximation to the standard errors of the parameter estimators via simulations. Additionally, we illustrate our example with a real-world linguistics dataset.