Study of Dimensionality Reduction and Clustering Machine Learning Algorithms for the Analysis of Ship Engine Data
摘要
Machine Learning (ML) is being successfully applied to ship engine management with proven economic and environmental benefits by engine performance optimization, timely fault detection and appropriate service planning. However, the data preparation for usage in ML algorithms provides several advantages including faster training and improved performance of the algorithm, improved visualization of the dataset, noise reduction, dataset simplification, avoidance of the curse of dimensionality and improved resource utilization. In this paper, two key techniques of the ML algorithms, that can be applied for data preparation and organization of ship engine data are studied, namely the dimensionality reduction and the data clustering. Dimensionality reduction involves the reduce of the number of input variables or features in a dataset, by retaining as much valuable information as possible. On the other hand, clustering ML techniques help to uncover insights and reduce data complexity through the organization of the data into clusters. Evaluation results demonstrate the usefulness of both techniques.