Digital transformation of ports is very important to deal with the problems of security, sustainability, and logistics. Precisely identifying various ship types is the aim of ship classification, which is essential for safeguarding interests and rights related to maritime trade and improving early warning systems for coastal defense. In this paper, an attempt is made to identify the impact of various feature selection methods like Chi-square, Correlation, Boruta, etc., on decision tree classifier performance in predicting the cargo code (NST2007) of Port of Sines, Portugal. The private dataset was collected by Port of Sines, Portugal for the period of June, 2023 to December, 2023. The model is trained on June dataset after pre-processing and tested on remaining months with same pre-processing steps. Various metrics are used to validate the model’s performance with various feature section methods. Finally, observed that model accuracy is degraded with the test cases due to non-seen cases which are not available during training the model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning for Cargo Code Prediction in Port Logistics

  • Karri Chiranjeevi,
  • Carlos Gomes,
  • Jose Moreira,
  • Luís Seabra Lopes,
  • Valentina Chkoniya,
  • Antonio Delgado Santos,
  • Sofia Martins,
  • Carla Marques

摘要

Digital transformation of ports is very important to deal with the problems of security, sustainability, and logistics. Precisely identifying various ship types is the aim of ship classification, which is essential for safeguarding interests and rights related to maritime trade and improving early warning systems for coastal defense. In this paper, an attempt is made to identify the impact of various feature selection methods like Chi-square, Correlation, Boruta, etc., on decision tree classifier performance in predicting the cargo code (NST2007) of Port of Sines, Portugal. The private dataset was collected by Port of Sines, Portugal for the period of June, 2023 to December, 2023. The model is trained on June dataset after pre-processing and tested on remaining months with same pre-processing steps. Various metrics are used to validate the model’s performance with various feature section methods. Finally, observed that model accuracy is degraded with the test cases due to non-seen cases which are not available during training the model.