<p>Machine learning algorithms have offered unprecedented solutions for many real-world problems. These algorithms frequently involve using a large number of features. However, several of these features could not be very informative due to data uncertainties, such as noise and residual variation. Decision trees are among the most preferred classification models. This is due to their simplicity, explainability, and readability. However, data inaccuracies could impact the construction of decision trees and thus hinder their results. Feature selection and construction present promising research direction to enhance the performance of decision tree models. In this paper, we present a strategy that combines feature selection and construction where the construction of new features is performed by using the ones that were not chosen during the selection step. However, the search space of combinations of selected/constructed features is extremely large. To find the best solution, a genetic algorithm has been developed combined with a graph covering vertices set guided approach. The obtained results on a large number of datasets from the UCI Repository demonstrate that our approach outperforms both recent and classical decision tree construction techniques. We also present a successful use case of our approach in detecting Botnet traffic in the Internet of Vehicles.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A genetic and graph-guided feature learning strategy for improving decision tree construction

  • Nour Elislem Karabadji,
  • Abdelaziz Amara Korba,
  • Ali Assi,
  • Hassina Seridi,
  • Mohamed Aimen Karabadji,
  • Yacine Ghamri-Doudane,
  • Abdelghani Lakhdari,
  • Mohamed Elati,
  • Wajdi Dhifli

摘要

Machine learning algorithms have offered unprecedented solutions for many real-world problems. These algorithms frequently involve using a large number of features. However, several of these features could not be very informative due to data uncertainties, such as noise and residual variation. Decision trees are among the most preferred classification models. This is due to their simplicity, explainability, and readability. However, data inaccuracies could impact the construction of decision trees and thus hinder their results. Feature selection and construction present promising research direction to enhance the performance of decision tree models. In this paper, we present a strategy that combines feature selection and construction where the construction of new features is performed by using the ones that were not chosen during the selection step. However, the search space of combinations of selected/constructed features is extremely large. To find the best solution, a genetic algorithm has been developed combined with a graph covering vertices set guided approach. The obtained results on a large number of datasets from the UCI Repository demonstrate that our approach outperforms both recent and classical decision tree construction techniques. We also present a successful use case of our approach in detecting Botnet traffic in the Internet of Vehicles.