<p>As the amount of unlabeled data has continued to grow and present challenges to machine learning practitioners, the need for unsupervised solutions is more evident than ever. With many unsupervised algorithms available to classify instances, the challenge remains that these algorithms require fine-tuning and/or appropriate parameter selection to produce reliable results. The difficulty remains that given an unlabeled dataset, the true class distribution is unknown, which impacts appropriateness of the selection of unsupervised algorithms and hyperparameter tuning, as well as the evaluation metrics chosen. Our novel approach addresses this critical gap in current literature. Through a fully automated and unsupervised framework, we take a binary unlabeled dataset, and return the class distribution without prior domain knowledge and regardless of the class distribution - imbalanced or balanced. We thoroughly investigate multiple datasets ranging in size, class distribution, and domain, and our empirical evidence demonstrates the successful determination of the class distribution given this variety of factors. Our approach uses data-driven threshold and parameter settings to improve model performance, particularly in imbalanced class scenarios. This helps in selecting suitable algorithms, guiding appropriate evaluation metrics, and promoting fairer, evidence-based decision making in fields such as fraud detection and healthcare.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel approach to automating unsupervised estimation of class distribution

  • Mary Anne Walauskis,
  • Taghi M. Khoshgoftaar

摘要

As the amount of unlabeled data has continued to grow and present challenges to machine learning practitioners, the need for unsupervised solutions is more evident than ever. With many unsupervised algorithms available to classify instances, the challenge remains that these algorithms require fine-tuning and/or appropriate parameter selection to produce reliable results. The difficulty remains that given an unlabeled dataset, the true class distribution is unknown, which impacts appropriateness of the selection of unsupervised algorithms and hyperparameter tuning, as well as the evaluation metrics chosen. Our novel approach addresses this critical gap in current literature. Through a fully automated and unsupervised framework, we take a binary unlabeled dataset, and return the class distribution without prior domain knowledge and regardless of the class distribution - imbalanced or balanced. We thoroughly investigate multiple datasets ranging in size, class distribution, and domain, and our empirical evidence demonstrates the successful determination of the class distribution given this variety of factors. Our approach uses data-driven threshold and parameter settings to improve model performance, particularly in imbalanced class scenarios. This helps in selecting suitable algorithms, guiding appropriate evaluation metrics, and promoting fairer, evidence-based decision making in fields such as fraud detection and healthcare.