Informativeness of Feature Sets in Data with Missing Values
摘要
The selection of informative sets of heterogeneous features is considered, taking into account the presence of gaps in the data when describing class objects. The selection condition is the property of data invariance to the scales of their measurements. Invariance is achieved through the use of methods for splitting feature values into non-intersecting intervals. Splitting into intervals is used in data preprocessing, carried out in order to unify the scales of feature measurements. Based on the results of preprocessing, sequences are formed that are ordered by the ratio of feature stability. It was studied the change in the order of features in sequences depending on the percentage of gaps in the data. Based on the percentage of gaps and the final conclusion, in the future, it is possible to assess the information content of various results of the experiments.