Sensor data streams are commonly used in many real-time Internet of Things (IoT) applications. However, some values are missing in the streams because of issues such as sensor malfunctions, intermittent communication errors, or drained batteries. These gaps in data can affect the accuracy of real-time analytics and other dependent processes. Current imputation methods either rely on predefined assumptions about observed streams or do not take advantage of the nature of data correlation, so these methods sometimes impute values inaccurately. It is for this reason that we aim to develop a more accurate and efficient imputation solution addressing missing values in data streams in order to ensure the operations of real-time applications. Firstly, we rely on a true assumption about the correlation between data points measured by sensors when they collect the same kinds of information. Secondly, we calculate the degree of correlation based on historical data of problematic sensors and others so as to eliminate data bias. After that, once the optimal correlative dataset is ready, we adapt and feed the continuous imputation framework MPIN with that dataset. Extensive experiments on a real environmental dataset show that our method achieves better results when taking data correlation into account compared to purely using the original one.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Missing Data Imputation for Sensor Observation Streams Leveraging Data Correlation and Message Propagation

  • Nguyen Thanh Quan,
  • Nguyen Ha Cong Ly

摘要

Sensor data streams are commonly used in many real-time Internet of Things (IoT) applications. However, some values are missing in the streams because of issues such as sensor malfunctions, intermittent communication errors, or drained batteries. These gaps in data can affect the accuracy of real-time analytics and other dependent processes. Current imputation methods either rely on predefined assumptions about observed streams or do not take advantage of the nature of data correlation, so these methods sometimes impute values inaccurately. It is for this reason that we aim to develop a more accurate and efficient imputation solution addressing missing values in data streams in order to ensure the operations of real-time applications. Firstly, we rely on a true assumption about the correlation between data points measured by sensors when they collect the same kinds of information. Secondly, we calculate the degree of correlation based on historical data of problematic sensors and others so as to eliminate data bias. After that, once the optimal correlative dataset is ready, we adapt and feed the continuous imputation framework MPIN with that dataset. Extensive experiments on a real environmental dataset show that our method achieves better results when taking data correlation into account compared to purely using the original one.