Corpora offer valuable, unbiased data for metaphor research. While methods exist for identifying metaphors in corpora, most of them focus on English. This chapter addresses this gap by establishing a method for identifying and annotating metaphors in a written Persian corpus, the first of its kind. We adapt the Metaphor Identification Procedure VU University Amsterdam (MIPVU) for Persian. This reliable method has been successfully applied to other languages with minor adjustments. Our goal is to provide a comprehensive explanation of this adaptation and construct the first Persian metaphor-annotated corpus. The process involves two steps: data preparation and data analysis. First, a 30,000-word corpus from online news articles (July 2023) was compiled from top Iranian agencies. Articles were randomly selected from common news categories. Part-of-speech tagging was performed using Stanford’s Stanza library. In the data analysis step, metaphorical annotations were assigned manually to content words. Then, inter-annotator agreement was calculated and descriptive statistics along with a qualitative analysis of the identified metaphors were presented. Python and Google Colab were used for coding. Finally, the annotated corpus was converted to XML format and made publicly available. This project establishes a foundation for building and annotating metaphor corpora in Persian. It contributes to metaphor identification research using corpora and showcases the successful adaptation of MIPVU for Persian, paving the way for similar endeavors in other languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards the First Persian Metaphor Annotated Corpus

  • Farzaneh Bakhtiyari,
  • Amirmasoud Iravani

摘要

Corpora offer valuable, unbiased data for metaphor research. While methods exist for identifying metaphors in corpora, most of them focus on English. This chapter addresses this gap by establishing a method for identifying and annotating metaphors in a written Persian corpus, the first of its kind. We adapt the Metaphor Identification Procedure VU University Amsterdam (MIPVU) for Persian. This reliable method has been successfully applied to other languages with minor adjustments. Our goal is to provide a comprehensive explanation of this adaptation and construct the first Persian metaphor-annotated corpus. The process involves two steps: data preparation and data analysis. First, a 30,000-word corpus from online news articles (July 2023) was compiled from top Iranian agencies. Articles were randomly selected from common news categories. Part-of-speech tagging was performed using Stanford’s Stanza library. In the data analysis step, metaphorical annotations were assigned manually to content words. Then, inter-annotator agreement was calculated and descriptive statistics along with a qualitative analysis of the identified metaphors were presented. Python and Google Colab were used for coding. Finally, the annotated corpus was converted to XML format and made publicly available. This project establishes a foundation for building and annotating metaphor corpora in Persian. It contributes to metaphor identification research using corpora and showcases the successful adaptation of MIPVU for Persian, paving the way for similar endeavors in other languages.