The Chinese grammar detection and correction system can automatically identify error positions and correct erroneous characters. Currently, there are still challenges in the research, such as the limited size of publicly available Chinese error correction datasets, the diversity of Chinese grammar error forms, and the difficulty in representing the distribution of grammar errors. Additionally, pre-trained Chinese language models lack the ability to differentiate between similar characters or words, affecting the accuracy of Chinese sentence detection and correction. In this paper, we propose a model based on grammar enhancement and feedback mechanism. During the model training phase, the Pairwise Character Interaction (PCI) module is used to enhance the grammatical representation of the text encoder. It employs various gating mechanisms based on character pairs to highlight the semantic features of grammatical errors at erroneous character positions. Furthermore, during the testing phase, our model (PCIFM) utilizes a feedback mechanism to edit and iteratively correct erroneous results. The proposed model was evaluated on publicly available CGED datasets, achieving the highest detection F1 scores on the test sets of CGED 2017, CGED 2018, and CGED 2020, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Chinese Grammar Correction Model Based on Semantic Enhancement and Feedback Mechanism

  • Zhujian Zhang,
  • Peiyu Zhao,
  • Bo Liu

摘要

The Chinese grammar detection and correction system can automatically identify error positions and correct erroneous characters. Currently, there are still challenges in the research, such as the limited size of publicly available Chinese error correction datasets, the diversity of Chinese grammar error forms, and the difficulty in representing the distribution of grammar errors. Additionally, pre-trained Chinese language models lack the ability to differentiate between similar characters or words, affecting the accuracy of Chinese sentence detection and correction. In this paper, we propose a model based on grammar enhancement and feedback mechanism. During the model training phase, the Pairwise Character Interaction (PCI) module is used to enhance the grammatical representation of the text encoder. It employs various gating mechanisms based on character pairs to highlight the semantic features of grammatical errors at erroneous character positions. Furthermore, during the testing phase, our model (PCIFM) utilizes a feedback mechanism to edit and iteratively correct erroneous results. The proposed model was evaluated on publicly available CGED datasets, achieving the highest detection F1 scores on the test sets of CGED 2017, CGED 2018, and CGED 2020, respectively.