Developing, Compiling and Annotating Corpora for the Persian Language
摘要
In this chapter, we briefly overview the general criteria that have to be taken into consideration while developing a corpus. Developing a corpus for the Persian language is challenging. In this chapter, the challenges are discussed and categorized. Then, we discuss the steps that have to be taken to make the developed corpus usable for natural language processing techniques. Furthermore, we explain how the data can be annotated. In the rest of this chapter, the sketch of supervised, semi-supervised, and unsupervised machine learning methods for data annotation is briefly explained. We collect 167 research papers that developed a corpus for their study; then, we categorize them based on the task and the corpus size. The annotated data has to be standardized. We briefly introduce the major standards used for structuring data.