错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Extraction of Morphological Categories for Morphemes

  • Abishek Stephen,
  • Vojtěch John,
  • Zdeněk Žabokrtský

摘要

Words in natural language can be assigned to specific morphological categories. For example, the English word ‘apples’ can be described using morphological labels like N;PL. The conditional probabilities on such word forms given the labels would reveal for English that the morpheme ‘s’ is present almost always when the label N;PL appears. This indicates that the morphological properties of a word can be traced to its morphemes. We do not have any data resource that associates morphemes with morphological categories. We use UniMorph schema and datasets for universal morphological annotation as a source of morphological categories and morpheme segmentation. We align morphemes (or exponents) with the corresponding morphological categories based on the UniMorph schema for 12 languages. Given the multilingual nature of the task, we utilize unsupervised methods based on the \(\Delta P\) measure and IBM Models as we test out the effectiveness of alignment methods used in statistical machine translation. Our results indicate that IBM Models accurately capture the alignment asymmetries between morphemes and morphological categories under non-trivial alignment settings.