Use of a Surrogate Model for Symbolic Discretization of Temporal Data Sets Through eMODiTS and a Training Set with Varying-Sized Instances
摘要
Time series classification is a supervised task in the field of temporal data mining. Time series naturally tend to be highly dimensional, requiring the use of reduction techniques such as discretization. eMODiTS is a data-driven method for symbolically discretizing time series, which determines the best scheme by modifying the number of time (word segments) and values (alphabet) cuts, generating a unique alphabet set for every word segment. However, due to the high computational cost required, a surrogate model is incorporated to minimize this cost, using the K-Nearest Neighbors approach for regression and Dynamic Time Warping (DTW) as the similarity measure. Results suggest that the surrogate model effectively estimates the objective functions’ values similarly to the original ones, leading to similar classification rates. It is validated with the statistical test where there is no significant statistical difference between the surrogate and original models. The surrogate model produces modified acceptance index ( \(d_j\) ) values regarding predicting ability, indicating that the predictive performance is on average. On the other hand, the Mean Squared Error (MSE) consistently stays below 0.15, demonstrating that even when surrogate models cannot estimate the same values as the original model, the similarity of the values remains clear.