<p>With the development of sensor observation technology, better data measurements at each component in engineering complex systems can be obtained. In this work, we consider and analyze the health state identification of complex experimental systems. Data-driven machine learning methods have become a suitable and timely approach in health state identification. However, limited by the environmental conditions and the algorithm, data-driven machine learning methods face the challenge of overfitting, robustness, and interpretability. We propose a method to select the feature variables for the training of the machine learning models based on the causal relationship of the measured variables. In this method, the fast causal inference technique is used to establish the causal relationships among the variables and then to get the causal graph. The eigenvector centrality and betweenness centrality of the causal graphs are shown to be key graph characteristics to find the criteria to select judiciously the smallest number of feature variables. Support vector machine (SVM) is used to predict the health state of two complex systems based on the selected feature variables. The two experiments results show that the feature set selected by our graph-based approach provides superior modelling accuracy than if the SVM is fed with variables selected by other methods in the literature, validating our theory on the feature selection for health state identification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Causal feature selection for health state identification of complex experimental systems

  • Chen Feng,
  • Murilo S. Baptista,
  • Xiaochen Liu,
  • Chenxing Ni,
  • Celso Grebogi

摘要

With the development of sensor observation technology, better data measurements at each component in engineering complex systems can be obtained. In this work, we consider and analyze the health state identification of complex experimental systems. Data-driven machine learning methods have become a suitable and timely approach in health state identification. However, limited by the environmental conditions and the algorithm, data-driven machine learning methods face the challenge of overfitting, robustness, and interpretability. We propose a method to select the feature variables for the training of the machine learning models based on the causal relationship of the measured variables. In this method, the fast causal inference technique is used to establish the causal relationships among the variables and then to get the causal graph. The eigenvector centrality and betweenness centrality of the causal graphs are shown to be key graph characteristics to find the criteria to select judiciously the smallest number of feature variables. Support vector machine (SVM) is used to predict the health state of two complex systems based on the selected feature variables. The two experiments results show that the feature set selected by our graph-based approach provides superior modelling accuracy than if the SVM is fed with variables selected by other methods in the literature, validating our theory on the feature selection for health state identification.