Increasing the Performance and Plausibility of Machine Learning via Data Analysis Techniques
摘要
The identification of significant input variables is crucial for the quality of machine learning models, since the models can only be as good as the representative quality of the data provided. In this chapter, different filter methods are compared as a subset of data analysis techniques for feature selection. Using real data from the power engineering field, the strengths of data analysis as a preprocessing step are demonstrated and the effectiveness of different filter methods is compared. The data analysis techniques not only increase the performance of machine learning models, but also their plausibility by identifying important input variables and reducing model complexity, resulting in more trustworthy and reliable models.