One of the most widely used corpora in the field of language teaching/learning is learner corpus, which mainly contains error annotations. The present research aimed at designing, developing, and annotating the first error-tagged Persian spoken learner corpus. The present corpus includes 10 h of data consisting of more than 71,000 words produced by 28 learners of Persian, which were manually transcribed and annotated. The total number of errors identified in the study was 5651, including 177 unique types of errors. The results showed that the highest number of errors at the level of appearance of error belonged to incorrect choice category, while the lowest number of errors was related to incorrect order. On the other hand, as far as the level of language category is concerned, the highest rate of errors was related to the category of syntax, whereas the lowest rate belonged to the lexis category. Finally, regarding the type of error level, vowel errors turned out to represent the highest error rate, while the lowest error rate belonged to the errors in intonation, comparative adjectives, and cardinal adjectives.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Design, Development, and Annotation of the Persian Spoken Learner Corpus

  • Ehsan Abdollahi,
  • Amirsaeid Moloodi,
  • Mohammad Rahimi,
  • Manouchehr Kouhestani

摘要

One of the most widely used corpora in the field of language teaching/learning is learner corpus, which mainly contains error annotations. The present research aimed at designing, developing, and annotating the first error-tagged Persian spoken learner corpus. The present corpus includes 10 h of data consisting of more than 71,000 words produced by 28 learners of Persian, which were manually transcribed and annotated. The total number of errors identified in the study was 5651, including 177 unique types of errors. The results showed that the highest number of errors at the level of appearance of error belonged to incorrect choice category, while the lowest number of errors was related to incorrect order. On the other hand, as far as the level of language category is concerned, the highest rate of errors was related to the category of syntax, whereas the lowest rate belonged to the lexis category. Finally, regarding the type of error level, vowel errors turned out to represent the highest error rate, while the lowest error rate belonged to the errors in intonation, comparative adjectives, and cardinal adjectives.