错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Data Distributions in Machine Learning Models with SOMs

  • Caroline König,
  • Alfredo Vellido

摘要

Data quality control is fundamental in data-driven analysis with machine learning (ML) models. In the domain of drug research, there is an increasing interest in the prediction of relevant biocompounds physicochemical properties with ML. In order to build predictive models of good quality, it is important to adequately select representative datasets. In this work, we combine ML prediction and Self-Organizing Maps-based exploration to build an interpretable machine learning model and to characterize those data that are most difficult to predict in the validation stage.