Relying on authentic data is a common prerequisite to almost all aspects of theoretical and applied linguistics research. This can be achieved by developing a large and representative corpus. Depending on the type of research, different specialized corpora have been compiled; among which are historical corpora These corpora consist of texts from one or more periods in the past. The corpus presented in this paper is the first attempt to build a corpus of Persian prose texts from the fifth to ninth centuries Anno Hegirae AH (tenth to fourteenth centuries Anno Domini) which consists of 70 authentic full texts with about five million words. The corpus is consisted of 1,039,734 types and 5,361,370 tokens, respectively. To compile this corpus, a collection of significant texts from this era was gathered according to specific criteria and formatted in a predetermined way. In the second step, all texts were integrated and indexed in the Persian Linguistic database (PLDB) system which facilitated text processing, faster information retrieval, and compiling frequency wordlists, statistics and concordances. The corpus is not just a raw ensemble of texts; the texts in PLDB are annotated with bibliographic headers, lemma, part of speech (POS), semantic (based on FARSNET) and phonetic tags.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Historical Corpus of the Persian Language

  • Saeideh Ghandi

摘要

Relying on authentic data is a common prerequisite to almost all aspects of theoretical and applied linguistics research. This can be achieved by developing a large and representative corpus. Depending on the type of research, different specialized corpora have been compiled; among which are historical corpora These corpora consist of texts from one or more periods in the past. The corpus presented in this paper is the first attempt to build a corpus of Persian prose texts from the fifth to ninth centuries Anno Hegirae AH (tenth to fourteenth centuries Anno Domini) which consists of 70 authentic full texts with about five million words. The corpus is consisted of 1,039,734 types and 5,361,370 tokens, respectively. To compile this corpus, a collection of significant texts from this era was gathered according to specific criteria and formatted in a predetermined way. In the second step, all texts were integrated and indexed in the Persian Linguistic database (PLDB) system which facilitated text processing, faster information retrieval, and compiling frequency wordlists, statistics and concordances. The corpus is not just a raw ensemble of texts; the texts in PLDB are annotated with bibliographic headers, lemma, part of speech (POS), semantic (based on FARSNET) and phonetic tags.