The extraction of one-dimensional (1D) vibrational data from structures as a means of continuous analysis has been a key focus of Structural Health Monitoring (SHM) for the past several decades. The recent advent of computer-aided algorithms in Machine Learning (ML) has allowed for the classification of 1D time series data without significant user intervention. Though these ML models are often robust concerning categorizing vibrational data, a few limitations are heightened by the constraints of the SHM domain. The most prevalent of these issues stems from the reliance of these models on massive, balanced, and complete datasets to train these models properly. Furthermore, the majority of the data collected often represents the “normal” condition of the structure, resulting in an imbalance between damaged and undamaged samples, resulting in models that are biased toward the class with the sample majority. Lastly, there are many instances where faulty sensors result in loss of data during instrumentation, further enforcing the data scarcity issues. To address these issues, numerous Data Augmentation (DA) techniques have been proposed, by which “synthetic” samples are created based on the characteristics of existing samples, which are then used to augment the existing limited, unbalanced dataset. However, few studies have investigated the correlation between the effectiveness of the DA technique with respect to the characteristics of the data and the methodology for analyzing the data. Therefore, this paper investigates the relationship between the effectiveness of DA techniques to data characteristics and model type. A 1D vibrational dataset is artificially imbalanced and rebalanced using data augmentation techniques including magnitude warping, permutation, and variational autoencoder. The performance of each “balanced” and “rebalanced” dataset for classification or prediction using various ML techniques such as 1D Convolutional Neural Networks, Support Vector Machines, and Random Forests is compared to characteristics of the data to determine the correlation between the DA chosen and final performance of the classification model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance Evaluation of Various Time Series Data Augmentation Techniques for Structural Vibration Application

  • Kyle Dunphy,
  • Zachary Baird,
  • Ayan Sadhu

摘要

The extraction of one-dimensional (1D) vibrational data from structures as a means of continuous analysis has been a key focus of Structural Health Monitoring (SHM) for the past several decades. The recent advent of computer-aided algorithms in Machine Learning (ML) has allowed for the classification of 1D time series data without significant user intervention. Though these ML models are often robust concerning categorizing vibrational data, a few limitations are heightened by the constraints of the SHM domain. The most prevalent of these issues stems from the reliance of these models on massive, balanced, and complete datasets to train these models properly. Furthermore, the majority of the data collected often represents the “normal” condition of the structure, resulting in an imbalance between damaged and undamaged samples, resulting in models that are biased toward the class with the sample majority. Lastly, there are many instances where faulty sensors result in loss of data during instrumentation, further enforcing the data scarcity issues. To address these issues, numerous Data Augmentation (DA) techniques have been proposed, by which “synthetic” samples are created based on the characteristics of existing samples, which are then used to augment the existing limited, unbalanced dataset. However, few studies have investigated the correlation between the effectiveness of the DA technique with respect to the characteristics of the data and the methodology for analyzing the data. Therefore, this paper investigates the relationship between the effectiveness of DA techniques to data characteristics and model type. A 1D vibrational dataset is artificially imbalanced and rebalanced using data augmentation techniques including magnitude warping, permutation, and variational autoencoder. The performance of each “balanced” and “rebalanced” dataset for classification or prediction using various ML techniques such as 1D Convolutional Neural Networks, Support Vector Machines, and Random Forests is compared to characteristics of the data to determine the correlation between the DA chosen and final performance of the classification model.