<p>Supervised deep learning methods for toponym recognition, which depend on substantial manually labeled data, often encounter challenges with the laborious task of annotation and difficulties in achieving model generalization. To address these issues, this paper introduces a semi-supervised toponym recognition approach that combines active learning and self-training techniques. First, the LDA topic model is used to classify the corpus data into diverse categories. Next, an active learning method based on the LTP uncertainty query strategy is deployed to further screen the samples. A curated selection of diverse and informative samples has been identified for manual labeling. Then, by combining the annotated samples with a large amount of unlabeled data, we employ a self-training method based on the BERT model that selects tokens with high confidence to iteratively update the model parameters, thereby enhancing the accuracy and generalization capabilities of our toponym recognition model. Finally, our method was applied to the MSRA, Ontonotes4, Boson, and People’s Daily datasets, with 1000, 800, 800, and 800 samples labeled respectively for experimentation. The comparative experimental results indicate that the F1 scores achieved on these four datasets are 0.9309, 0.7732, 0.7914, and 0.8996, respectively, which are higher than those of the semi-supervised baseline models. It is proved that the method proposed in this paper achieves a good performance of toponym recognition with a small amount of labeling data, significantly reducing the workload of manual labeling.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A semi-supervised Chinese toponym recognition methods combining active learning and self-training

  • Zhao Yijiang,
  • Luo Jing,
  • Zhang Daoan,
  • Liu Yizhi,
  • Liao Zhuhua

摘要

Supervised deep learning methods for toponym recognition, which depend on substantial manually labeled data, often encounter challenges with the laborious task of annotation and difficulties in achieving model generalization. To address these issues, this paper introduces a semi-supervised toponym recognition approach that combines active learning and self-training techniques. First, the LDA topic model is used to classify the corpus data into diverse categories. Next, an active learning method based on the LTP uncertainty query strategy is deployed to further screen the samples. A curated selection of diverse and informative samples has been identified for manual labeling. Then, by combining the annotated samples with a large amount of unlabeled data, we employ a self-training method based on the BERT model that selects tokens with high confidence to iteratively update the model parameters, thereby enhancing the accuracy and generalization capabilities of our toponym recognition model. Finally, our method was applied to the MSRA, Ontonotes4, Boson, and People’s Daily datasets, with 1000, 800, 800, and 800 samples labeled respectively for experimentation. The comparative experimental results indicate that the F1 scores achieved on these four datasets are 0.9309, 0.7732, 0.7914, and 0.8996, respectively, which are higher than those of the semi-supervised baseline models. It is proved that the method proposed in this paper achieves a good performance of toponym recognition with a small amount of labeling data, significantly reducing the workload of manual labeling.