In coal mine safety research, the precise classification of accident reports is paramount. This process facilitates a rapid comprehension of accident causes, enabling the formulation of effective preventive measures. The advent of Natural Language Processing (NLP), particularly through the emergence of models like BERT and its variants, has revolutionized our capacity for accurate report classification. Yet, the challenge of scarce coal mine safety labeled data and the prohibitive costs of labeling persists, impeding the optimal utilization of pre-trained models. Data augmentation proves to be a powerful method for addressing these challenges. However, traditional text data augmentation techniques face limitations due to their potential to generate homogenous data and the risk of distorting essential information. To overcome these constraints, we introduce ”AugMine,” an innovative text data augmentation strategy. AugMine capitalizes on ChatGPT’s prowess in generating high-quality text from limited datasets, thereby broadening the training pool and bolstering the model’s proficiency in identifying critical components within coal mine accident reports. Furthermore, we incorporate adversarial training techniques to further enhance classification performance. In this study, we leveraged BERT and its derivatives for feature extraction and assessed multiple data augmentation strategies. The experimental results demonstrate that the “AugMine” approach, which we introduced, notably enhances the precision of classifying coal mine accident reports, outperforming established text data augmentation techniques.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AugMine: Boosting Coal Mine Accident News Classification with Text Data Augmentation

  • He Hu,
  • Yixu Feng,
  • Chaoqun Wang,
  • Zhaohe Wang,
  • Xiaowen Ma,
  • Peng Wu,
  • Wei Dong,
  • Qingsen Yan

摘要

In coal mine safety research, the precise classification of accident reports is paramount. This process facilitates a rapid comprehension of accident causes, enabling the formulation of effective preventive measures. The advent of Natural Language Processing (NLP), particularly through the emergence of models like BERT and its variants, has revolutionized our capacity for accurate report classification. Yet, the challenge of scarce coal mine safety labeled data and the prohibitive costs of labeling persists, impeding the optimal utilization of pre-trained models. Data augmentation proves to be a powerful method for addressing these challenges. However, traditional text data augmentation techniques face limitations due to their potential to generate homogenous data and the risk of distorting essential information. To overcome these constraints, we introduce ”AugMine,” an innovative text data augmentation strategy. AugMine capitalizes on ChatGPT’s prowess in generating high-quality text from limited datasets, thereby broadening the training pool and bolstering the model’s proficiency in identifying critical components within coal mine accident reports. Furthermore, we incorporate adversarial training techniques to further enhance classification performance. In this study, we leveraged BERT and its derivatives for feature extraction and assessed multiple data augmentation strategies. The experimental results demonstrate that the “AugMine” approach, which we introduced, notably enhances the precision of classifying coal mine accident reports, outperforming established text data augmentation techniques.