The construction projects’ data collected using IoT devices is characterized by its high velocity as being streamed continuously and its varying accuracy and reliability due to possible sensor malfunctions, transmission errors and other factors, so AI technologies-based system must be able to preprocess the collected data. The use of the wrapper technique to filter out irrelevant and redundant information has two associated problems: a high computation costs and a risk to receive an overfitting learning model. In this paper the focus is put on a design of a mathematical model to be used with a search strategy in a wrapper technique to help to reduce a computation cost and to prevent a classification machine learning model to overfit. To meet the goal in the research were used the methods: data analysis and visualization; mathematical modelling, structuring by algorithm; experimental tests. The results of the experiments are analyzed and in the conclusions is recommended to use this study findings to improve the classification machine learning strategy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature Search Strategy for Wrapper Techniques to Prevent Overfitting in Classification Models

  • Olga Solovei,
  • Tetyana Honcharenko

摘要

The construction projects’ data collected using IoT devices is characterized by its high velocity as being streamed continuously and its varying accuracy and reliability due to possible sensor malfunctions, transmission errors and other factors, so AI technologies-based system must be able to preprocess the collected data. The use of the wrapper technique to filter out irrelevant and redundant information has two associated problems: a high computation costs and a risk to receive an overfitting learning model. In this paper the focus is put on a design of a mathematical model to be used with a search strategy in a wrapper technique to help to reduce a computation cost and to prevent a classification machine learning model to overfit. To meet the goal in the research were used the methods: data analysis and visualization; mathematical modelling, structuring by algorithm; experimental tests. The results of the experiments are analyzed and in the conclusions is recommended to use this study findings to improve the classification machine learning strategy.