<p>Data stream classification is an important research direction in the field of data mining. In the early stages of the emergence of new concepts in data streams, the issue of label scarcity limits the effectiveness of supervised learning. Meanwhile, evolutionary phenomena in data streams significantly affect classification performance. To address the issue of label scarcity in evolving data streams, this paper proposes an online transfer learning framework, referred to as OTLF. To reduce the distribution discrepancy between the source domain and the initial target domain and enhance transferability, OTLF introduces a clustering-based data preprocessing method for the source domain. This method selects the source instances that are most similar to the target domain distribution by calculating similarity weights. To improve the adaptability of the model to data stream evolution and enhance transfer efficiency, OTLF first updates and maintains the source domain micro-clusters online to effectively capture concept drift. It then constructs a concept evolution detection module, which uses a buffer to store potential new class anomalies and employs an emerging class detection method to monitor the appearance of new classes. Comparative simulation experiments with other algorithms on different data streams show that OTLF is effective and stable in most cases. Furthermore, OTLF maintains a competitive advantage in terms of classification accuracy and other performance metrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Online transfer learning framework for label scarcity in evolving data streams

  • Jian Zhu,
  • Sanmin Liu,
  • Subin Huang,
  • Ping Zhang,
  • Wentao Yu,
  • Tuyi Zhang,
  • Guoyi Zhang

摘要

Data stream classification is an important research direction in the field of data mining. In the early stages of the emergence of new concepts in data streams, the issue of label scarcity limits the effectiveness of supervised learning. Meanwhile, evolutionary phenomena in data streams significantly affect classification performance. To address the issue of label scarcity in evolving data streams, this paper proposes an online transfer learning framework, referred to as OTLF. To reduce the distribution discrepancy between the source domain and the initial target domain and enhance transferability, OTLF introduces a clustering-based data preprocessing method for the source domain. This method selects the source instances that are most similar to the target domain distribution by calculating similarity weights. To improve the adaptability of the model to data stream evolution and enhance transfer efficiency, OTLF first updates and maintains the source domain micro-clusters online to effectively capture concept drift. It then constructs a concept evolution detection module, which uses a buffer to store potential new class anomalies and employs an emerging class detection method to monitor the appearance of new classes. Comparative simulation experiments with other algorithms on different data streams show that OTLF is effective and stable in most cases. Furthermore, OTLF maintains a competitive advantage in terms of classification accuracy and other performance metrics.