In this chapter, we discuss and apply a regression-based data imputation method to a real-world dataset when the data are a sequence (time series) of vectors (multidimensional measurements). Due to the significant size of the measurement vector in relation to the number of observed vectors, the PCA-based method of dimension reduction is applied. Based on incomplete data vectors, a neural network is created, the output of which are the values of the coefficients in the model based on the principal components. The values of these coefficients obtained from the net lead to an approximate reconstruction of the entire data vector (observation vector in which the values of specific coordinates are missing). Due to the nature of the data, our imputation model is a simulation model and not a prediction model. The experiments presented in the chapter were carried out with real data that were the sequence of the daily average levels of PM10 dust pollution at the 374 measurement points located in Poland.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Use of Neural Networks and PCA Dimensionality Reduction in the Imputation of Missing Fragments in High-Dimensional Time Series

  • Ewa Skubalska-Rafajlowicz,
  • Adam Krzyżak,
  • Michal Piórek

摘要

In this chapter, we discuss and apply a regression-based data imputation method to a real-world dataset when the data are a sequence (time series) of vectors (multidimensional measurements). Due to the significant size of the measurement vector in relation to the number of observed vectors, the PCA-based method of dimension reduction is applied. Based on incomplete data vectors, a neural network is created, the output of which are the values of the coefficients in the model based on the principal components. The values of these coefficients obtained from the net lead to an approximate reconstruction of the entire data vector (observation vector in which the values of specific coordinates are missing). Due to the nature of the data, our imputation model is a simulation model and not a prediction model. The experiments presented in the chapter were carried out with real data that were the sequence of the daily average levels of PM10 dust pollution at the 374 measurement points located in Poland.