<p>Concept drift (and data drift) is a common phenomenon in machine learning models, where the statistical properties of the input data change over time, leading to a decrease in model performance. Detecting data drift is crucial for maintaining the accuracy and reliability of machine learning models in real-world applications. While previous data drift detector approaches can identify if a drift has occurred, these approaches cannot localize which specific features have caused the drift. Feature drift detectors solve this deficiency, but the required number of detectors is equal to the number of dimensions, which is a resource-intensive solution in high-dimensional data. In this paper, we propose a novel approach for feature drift analysis and drift detection based on a domino effect caused by the correlation of features. Our approach, the so-called Domino drift effect (DDE), is based on the empirically proven assumption that an initial reference correlation can be utilized as a proxy for detecting other drifting features. The method analyzes the correlating and drifting behavior, and by using only a subset of all features, it derives inference about the drifting of the remaining features, if co-drifting phenomena occur in the data stream. At co-drifting phenomena, the DDE method can estimate the probability of feature drift, which is particularly useful in high-dimensional datasets. To evaluate the effectiveness of our approach, we conducted experiments on four real-world datasets. The results show that our approach can effectively be used to predict feature drift in the whole dataset, and it has potential industrial applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Domino drift effect approach for probability estimation of feature drift in high-dimensional data

  • Gábor Szűcs,
  • Marcell Németh

摘要

Concept drift (and data drift) is a common phenomenon in machine learning models, where the statistical properties of the input data change over time, leading to a decrease in model performance. Detecting data drift is crucial for maintaining the accuracy and reliability of machine learning models in real-world applications. While previous data drift detector approaches can identify if a drift has occurred, these approaches cannot localize which specific features have caused the drift. Feature drift detectors solve this deficiency, but the required number of detectors is equal to the number of dimensions, which is a resource-intensive solution in high-dimensional data. In this paper, we propose a novel approach for feature drift analysis and drift detection based on a domino effect caused by the correlation of features. Our approach, the so-called Domino drift effect (DDE), is based on the empirically proven assumption that an initial reference correlation can be utilized as a proxy for detecting other drifting features. The method analyzes the correlating and drifting behavior, and by using only a subset of all features, it derives inference about the drifting of the remaining features, if co-drifting phenomena occur in the data stream. At co-drifting phenomena, the DDE method can estimate the probability of feature drift, which is particularly useful in high-dimensional datasets. To evaluate the effectiveness of our approach, we conducted experiments on four real-world datasets. The results show that our approach can effectively be used to predict feature drift in the whole dataset, and it has potential industrial applications.