错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SSDIR: A Novel Semi-supervised Approach for Data Imbalanced Regression

  • Hubo Yuan,
  • Shuo Wang

摘要

Data imbalance persists as a critical challenge across real-world applications, causing models to underperform on vital minority patterns. While semi-supervised learning has been used to facilitate imbalanced classification, imbalanced regression remains underexplored due to continuous output value. We propose a self-training framework that makes use of unlabeled data to tackle imbalanced regression problems, called SSDIR. It introduces two key innovations: 1) a feature-aware label confidence algorithm that evaluates sample reliability through importance-guided perturbation and iterative verification, and 2) a dynamic category-aware sampling strategy optimized for continuous label distributions. Comprehensive experiments are conducted on the real-world datasets to validate our framework’s effectiveness. These datasets, representing diverse regression scenarios with inherent output imbalance, demonstrate SSDIR’s consistent improvement by 61% in capturing minority patterns while maintaining overall performance. The results provide empirical foundations for advancing imbalanced regression research.