Discrimination of natural and nonnatural earthquakes using a vision transformer
摘要
Rapidly and reliably distinguishing between natural and nonnatural small-scale earthquakes is crucial for earthquake monitoring and seismic activity analyses. In this study, we propose a vision transformer for seismology (SeisViT) to discriminate between natural and non-natural earthquakes. Our SeisViT is based on a vision transformer (ViT) network that introduces a multihead self-attention mechanism, which can effectively capture and focus on important features from seismic waveforms.The SeisViT model processes three-component raw waveforms from a single seismic station, using data collected from natural and nonnatural earthquakes in China. Through a comprehensive evaluation of hyperparameters—including learning rate, number of transformer encoder layers, and patch size-we optimized the SeisViT architecture to achieve maximal performance. Our results demonstrate that the SeisViT model, with a learning rate of 10-3, six transformer encoder layers, and a patch size of eight, achieves superior accuracy in discriminating natural from nonnatural earthquakes. Compared to conventional models such as multilayer perceptron (MLP), decision tree (DT), random forest (RF), and support vector machine (SVM), the SeisViT model achieved the highest accuracy (90.17%), precision (89.68%), recall (89.90%), and F1 score (89.79%) on the test dataset. These results underscore the potential of the SeisViT model as a significant advancement for earthquake monitoring with promising applications in seismology.