In the realm of deep learning, the efficacy of models is inextricably linked to the quality of the training datasets. This relationship is particularly pronounced when the data in question exhibits complex and sparse characteristics, as is often the case with electromagnetic wave data. The conventional visual assessment of dataset quality becomes inadequate due to the intangible nature of these features. Addressing this challenge, this paper presents a pioneering method for the evaluation of dataset quality. Our approach draws upon high-dimensional feature data extracted via a backbone neural network, and it is influenced by the principles of the Image Structure Similarity (SSIM) index. This innovative technique facilitates the quantification and scoring of dataset quality for arbitrary batches of feature sets. Empirical evidence suggests that the utilization of datasets with higher scores, as determined by our method, significantly enhances the performance of deep learning models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Dataset Diversity Evaluation Method via High-Dimensional Feature Representation

  • Zichao Zhang,
  • Junjie Ma,
  • Weiyu Ji,
  • Landi Gu,
  • Wei Jiang

摘要

In the realm of deep learning, the efficacy of models is inextricably linked to the quality of the training datasets. This relationship is particularly pronounced when the data in question exhibits complex and sparse characteristics, as is often the case with electromagnetic wave data. The conventional visual assessment of dataset quality becomes inadequate due to the intangible nature of these features. Addressing this challenge, this paper presents a pioneering method for the evaluation of dataset quality. Our approach draws upon high-dimensional feature data extracted via a backbone neural network, and it is influenced by the principles of the Image Structure Similarity (SSIM) index. This innovative technique facilitates the quantification and scoring of dataset quality for arbitrary batches of feature sets. Empirical evidence suggests that the utilization of datasets with higher scores, as determined by our method, significantly enhances the performance of deep learning models.