Discretisation and Attribute Relevance in Knowledge Mining Problems
摘要
The chapter presents research dedicated to the relevance of attributes in machine learning tasks, observed from various perspectives. Firstly, information on the importance of features can come from expert domain knowledge. Secondly, there exist numerous measures and algorithms that can be applied to the data and which return some estimation of relevance. Also, in the case of continuous variables, the methods employed for their pre-processing and used for supervised discretisation can detect how much the available features support recognition of classes, which can be treated to some extent as insight into their importance for predictions. When attributes are translated into their discrete domains, the resulting variants of data can be analysed to study how such transformations changed the relevance by evaluating with the algorithms used previously in the continuous domain. Furthermore, the influence of discretisation on patterns existing in the data can be examined from the perspective of performance of selected inducers, capable of operating both on real-valued and categorical attributes. The results of the research works indicate that the discretisation procedure can noticeably affect the importance of features, in particular when extended discretisation procedures are applied, combining supervised and unsupervised methods. Such processing makes the best of all available features and can result in increased importance for some attributes, at the same time enhancing performance.