This paper presents a new dataset of noun-adjective multiword expressions with different degrees of compositionality and semantic ambiguity in Galician. It is composed of 240 MWEs, which can convey one or different senses depending on the context. For each sense, a language expert manually created two sentences and selected from corpora four additional examples that included the target MWEs, thus resulting in a useful resource for exploring potential data contamination when evaluating language models. Each MWE in context was then classified as idiomatic, partially idiomatic, or compositional. Therefore, the dataset comprises MWEs with stable meanings, and two types of ambiguous expressions: 1) potential idiomatic expressions (e.g., red flag), and 2) polysemy-based ambiguous MWEs, whose various senses are due to the ambiguity of one of the constituent words (e.g., common noun as a type of noun or a noun that is common). To illustrate the potential of this resource, a comparison of three BERT models for Galician was performed, shedding light on the representation of ambiguous MWEs in Transformers. This is a valuable resource for evaluating the semantic capabilities of current language models in a low-resource variety, bearing in mind that idiomaticity is one of the linguistic phenomena whose modeling poses the greatest challenges for computational approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Compositionality and Ambiguity in Multiword Expressions: A Dataset for the Evaluation of Language Models in Galician

  • Laura Castro,
  • Anna Temerko,
  • Marcos Garcia

摘要

This paper presents a new dataset of noun-adjective multiword expressions with different degrees of compositionality and semantic ambiguity in Galician. It is composed of 240 MWEs, which can convey one or different senses depending on the context. For each sense, a language expert manually created two sentences and selected from corpora four additional examples that included the target MWEs, thus resulting in a useful resource for exploring potential data contamination when evaluating language models. Each MWE in context was then classified as idiomatic, partially idiomatic, or compositional. Therefore, the dataset comprises MWEs with stable meanings, and two types of ambiguous expressions: 1) potential idiomatic expressions (e.g., red flag), and 2) polysemy-based ambiguous MWEs, whose various senses are due to the ambiguity of one of the constituent words (e.g., common noun as a type of noun or a noun that is common). To illustrate the potential of this resource, a comparison of three BERT models for Galician was performed, shedding light on the representation of ambiguous MWEs in Transformers. This is a valuable resource for evaluating the semantic capabilities of current language models in a low-resource variety, bearing in mind that idiomaticity is one of the linguistic phenomena whose modeling poses the greatest challenges for computational approaches.