Introducing lexicon knowledge into character-level models enhances their ability to discern word boundaries, thus boosting the model’s performance. Although the above methods have significantly improved the performance of the model, there are still the following issues. The existing character-level models can only obtain word boundary information from potential lexicon generated from text, and special word information that does not appear in the lexicon cannot be perceived and obtained. Therefore, the problem of out-of-vocabulary words in the lexicon will reduce the performance of the model. To address this, we propose a Chinese Named Entity Recognition Based on Template and Contrastive Learning model(TBCL). This model segments input sentences into n-gram spans and constructs prompt templates to capture comprehensive word information. Leveraging contrastive learning, we designed a triplet loss function to efficiently utilize prompt templates for learning word boundary information. The model is tested on three Chinese named entity recognition datasets: Resume, MSRA and OntoNotes, and the experiment results obtained are superior to the baseline models, which fully prove the effectiveness of the proposed model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Chinese Named Entity Recognition Based on Template and Contrastive Learning

  • Jingjing Zhu,
  • Tianyu Cai,
  • Zhenyu Zhao,
  • Shenggen Ju

摘要

Introducing lexicon knowledge into character-level models enhances their ability to discern word boundaries, thus boosting the model’s performance. Although the above methods have significantly improved the performance of the model, there are still the following issues. The existing character-level models can only obtain word boundary information from potential lexicon generated from text, and special word information that does not appear in the lexicon cannot be perceived and obtained. Therefore, the problem of out-of-vocabulary words in the lexicon will reduce the performance of the model. To address this, we propose a Chinese Named Entity Recognition Based on Template and Contrastive Learning model(TBCL). This model segments input sentences into n-gram spans and constructs prompt templates to capture comprehensive word information. Leveraging contrastive learning, we designed a triplet loss function to efficiently utilize prompt templates for learning word boundary information. The model is tested on three Chinese named entity recognition datasets: Resume, MSRA and OntoNotes, and the experiment results obtained are superior to the baseline models, which fully prove the effectiveness of the proposed model.