错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Closing the Data Gap: A Comparative Study of Missing Value Imputation Algorithms in Time Series Datasets

  • Sepideh Hassankhani Dolatabadi,
  • Ivana Budinská,
  • Rafe Behmaneshpour,
  • Emil Gatial

摘要

The presence of missing values in time series datasets poses significant challenges for accurate data analysis and modeling. In this paper, we present a comparative study of missing value imputation algorithms applied to time series datasets collected from various sensors over a period of six months. The goal of this study is to bridge the data gap by effectively replacing missing values and assessing the performance of three common imputation algorithms for time series: K-Nearest Neighbors (KNN) imputer, Expectation-Maximization (EM), and Multiple Imputation by Chained Equations (MICE). To evaluate the performance of the imputation techniques, we employed Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE) as metrics. Through rigorous experimentation and analysis, we found that each algorithm exhibited varying degrees of effectiveness in handling missing values within the time series datasets. Our findings highlight the importance of choosing an appropriate imputation algorithm based on the characteristics of the dataset and the specific requirements of the analysis. The results also demonstrate the potential of the MICE imputer in closing the data gap and improving the accuracy of subsequent analyses on time series sensor data. Overall, this study provides valuable insights into the performance and suitability of different missing value imputation algorithms for time series datasets, facilitating better decision-making and enhancing the reliability of data-driven applications in various domains.