Performance Evaluation of Classification Model in Automotive Industry with Machine Learning
摘要
The automotive industry is undergoing an innovative transformation with machine learning, artificial intelligence, and big data. The integration of software into cars provides a safer driving experience for vehicle drivers, which has started to attract consumer interest in software usage in the automotive sector. Machine learning offers the potential to reduce costs, increase efficiency, predict prices, enhance customer satisfaction, and gain a competitive advantage. Therefore, machine learning is becoming increasingly important for automotive companies. Furthermore, companies in the automotive sector continue to focus on these new technologies to gain a competitive edge over their competitors. Machine learning studies in this field contribute to accurate pricing by buyers and sellers, help manufacturers adapt to consumer preferences, and provide significant advantages for the automotive industry and consumers. In this thesis study conducted on the automobile information dataset, automobile classification was performed using logistic regression and random forest classification algorithms among Europe, Japan, and the United States (US). The performance of the classification models was evaluated based on the results obtained, and a decision was made on which model to use. According to the evaluated performance metrics, the random forest model outperforms the logistic regression model. It has higher accuracy rates and macro and weighted averages. When the confusion matrix results are examined, it is observed that the second model has higher accuracy rates, especially in the Europe and US classes, and is acceptable in the Japan class as well. The random forest model generally has higher accuracy and better macro and weighted averages. The accuracy rate of the logistic regression model is 70%, while that of the random forest model is 82%, and it demonstrates a more balanced performance among classes. Consequently, according to the performance evaluation study of the classification model conducted with this dataset, the random forest model should be selected.