Towards High-Performance Exploratory Data Analysis (EDA) via Stable Equilibrium Point
摘要
Exploratory data analysis (EDA) is a vital procedure in data science projects. In this work, we introduce a stable equilibrium point (SEP)-based framework for improving the performance of EDA. By exploiting the SEPs to be the representative points, our approach aims to generate high-quality clustering and data visualization for real-world data sets. A very unique property of the proposed method is that the SEPs will directly encode the clustering properties of data sets. Compared with prior state-of-the-art clustering and data visualization methods, the proposed methods allow substantially improving solution quality for large-scale data analysis tasks. For instance, for the USPS data set, our method achieves more than \(10\%\) clustering accuracy gain over the standard spectral clustering algorithm and 3X speedup for the t-SNE visualization.