With the growing demand for computility, the reliability of computility services has become increasingly crucial. Due to the escalating volume and complexity of tasks processed, computility services often need to operate under high load, which can easily lead to issues such as resource shortages and service interruptions. Logs in computility services meticulously record the operational information of each component; therefore, anomaly detection based on logs can effectively ensure the stable operation of computility services. This study aims to address two challenges in the field of log anomaly detection. First, this study addresses the previously overlooked issue of class-imbalanced log data. Second, given the massive volumes of log data, the time required for model training poses a significant challenge. To address these issues, we propose EDSLog, a novel efficient log anomaly detection framework based on dataset partitioning. Initially, EDSLog processes log sequences through the Weight-Based K-fold Sub Hold-out Method (WKHM), effectively alleviating the class-imbalance problem. Subsequently, EDSLog leverages Simple Recurrent Units (SRU) enhanced by a self-attention mechanism to extract features from log sequences. Finally, EDSLog determines whether the predicted log data are anomalous. Experiments show that EDSLog achieves the best evaluation metrics in class-imbalanced datasets while having the shortest total model runtime. Specifically, EDSLog achieved the highest F1 scores of 100 and 99.96 respectively on the BGL and HDFS datasets, where abnormal logs account for 0.1% of the data. Additionally, EDSLog’s training speed was 35.62% faster than the model with the second shortest training duration among all models compared.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EDSLog: Efficient Log Anomaly Detection Method Based on Dataset Partitioning

  • Feng Liang,
  • Jing Liu

摘要

With the growing demand for computility, the reliability of computility services has become increasingly crucial. Due to the escalating volume and complexity of tasks processed, computility services often need to operate under high load, which can easily lead to issues such as resource shortages and service interruptions. Logs in computility services meticulously record the operational information of each component; therefore, anomaly detection based on logs can effectively ensure the stable operation of computility services. This study aims to address two challenges in the field of log anomaly detection. First, this study addresses the previously overlooked issue of class-imbalanced log data. Second, given the massive volumes of log data, the time required for model training poses a significant challenge. To address these issues, we propose EDSLog, a novel efficient log anomaly detection framework based on dataset partitioning. Initially, EDSLog processes log sequences through the Weight-Based K-fold Sub Hold-out Method (WKHM), effectively alleviating the class-imbalance problem. Subsequently, EDSLog leverages Simple Recurrent Units (SRU) enhanced by a self-attention mechanism to extract features from log sequences. Finally, EDSLog determines whether the predicted log data are anomalous. Experiments show that EDSLog achieves the best evaluation metrics in class-imbalanced datasets while having the shortest total model runtime. Specifically, EDSLog achieved the highest F1 scores of 100 and 99.96 respectively on the BGL and HDFS datasets, where abnormal logs account for 0.1% of the data. Additionally, EDSLog’s training speed was 35.62% faster than the model with the second shortest training duration among all models compared.