Multi-label classifiers make use of associations between labels in multi-label data to increase the accuracy of prediction. Before using a multi-label classifier, the data should be analysed to identify if there are associations between labels. If sets of independent labels are found, the data can be split in to multiple smaller data sets for analysis. Unfortunately, each label is dependent on the set of observations and so measuring label dependence is futile. What we actually seek is independence after taking the observations into account. In this article, we examine the concepts of explained and unexplained label covariance for measuring label dependence. We explore the use of a Normal copula model for modelling the label dependence/covariance and show that it is not able to measure conditional covariance directly. We then propose a new statistical model that allows direct measurement of label covariance (both constant and conditional). The model is validated using generated data and it is also used to examine the label covariance in real world data, allowing us to build simpler multi-label models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Determining the Need for Multi-label Classifiers by Measuring Unexplained Covariance

  • Laurence A. F. Park,
  • Jesse Read

摘要

Multi-label classifiers make use of associations between labels in multi-label data to increase the accuracy of prediction. Before using a multi-label classifier, the data should be analysed to identify if there are associations between labels. If sets of independent labels are found, the data can be split in to multiple smaller data sets for analysis. Unfortunately, each label is dependent on the set of observations and so measuring label dependence is futile. What we actually seek is independence after taking the observations into account. In this article, we examine the concepts of explained and unexplained label covariance for measuring label dependence. We explore the use of a Normal copula model for modelling the label dependence/covariance and show that it is not able to measure conditional covariance directly. We then propose a new statistical model that allows direct measurement of label covariance (both constant and conditional). The model is validated using generated data and it is also used to examine the label covariance in real world data, allowing us to build simpler multi-label models.