The present paper introduces a multilingual data set of erroneous and correct text sentences. The novel data set marks a significant advancement from an existing corpus by incorporating additional samples and refining its overall structure. The primary purpose of this data set is to support the research and development of automated error detection systems, especially in the multilingual setting where high-quality data sets are scarce. A distinctive feature of our data set is that it incorporates only incorrect sentences and their corresponding correct versions. These sentences are sourced from a variety of texts written by native speakers from different industries, such as pharmaceuticals, banking, insurance, retail, communications, and more. Each sentence in the data set has been annotated by professional proofreaders. The paper includes a comprehensive error analysis, where we classify and scrutinize the different types of errors within the data set. By categorizing and analysing the errors in the data set, we aim to identify patterns and common issues. Additionally, we conduct a thorough experimental evaluation using a well-established language model. Our analysis assesses the classification accuracy measured over all errors and the accuracy of each specific error type. Interestingly, our results show that while some error types can be detected with an accuracy exceeding 80%, it turns out that the recognition of other error types is very difficult to solve.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Error Detection Through Specialized Task Implementation

  • Corina Masanti,
  • Hans-Friedrich Witschel,
  • Kaspar Riesen

摘要

The present paper introduces a multilingual data set of erroneous and correct text sentences. The novel data set marks a significant advancement from an existing corpus by incorporating additional samples and refining its overall structure. The primary purpose of this data set is to support the research and development of automated error detection systems, especially in the multilingual setting where high-quality data sets are scarce. A distinctive feature of our data set is that it incorporates only incorrect sentences and their corresponding correct versions. These sentences are sourced from a variety of texts written by native speakers from different industries, such as pharmaceuticals, banking, insurance, retail, communications, and more. Each sentence in the data set has been annotated by professional proofreaders. The paper includes a comprehensive error analysis, where we classify and scrutinize the different types of errors within the data set. By categorizing and analysing the errors in the data set, we aim to identify patterns and common issues. Additionally, we conduct a thorough experimental evaluation using a well-established language model. Our analysis assesses the classification accuracy measured over all errors and the accuracy of each specific error type. Interestingly, our results show that while some error types can be detected with an accuracy exceeding 80%, it turns out that the recognition of other error types is very difficult to solve.