A Method for Comparative Visualization of Labeled Multidimensional Data and Its Application to Machine Learning Data
摘要
This chapter proposes a method for visualizing differences among labeled multidimensional data. The proposed method arranges multiple given multidimensional datasets on the same screen space applying the same dimensionality reduction scheme, and then displays a group of samples semi-transparently representing the labels with their particular colors. This representation makes it easy to observe which labels have the most in common or differences among the multidimensional data, and which labels tend to cause outliers. This chapter presents an application example using MNIST, USPS, and CIFAR10, which are representative sample datasets for machine learning, and discusses the effectiveness and issues of the proposed method.