错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Needle in a Haystack: Finding Suitable Idioms Based on Text Descriptions

  • Dmitrii Zhernokleev,
  • Pavel Braslavski

摘要

Idioms are an important part of natural languages and are often used to enhance expressiveness and fluency of speech. However, it can be difficult to find a contextually appropriate idiom when writing an essay or crafting a headline for a news article, especially for non-native speakers. This gives rise to the idea of an automated system that is able to recommend an idiom for an input sentence. The goal of this study is to develop and compare methods that would make such a system possible. We used an existing collection of idioms and employed several configurations based on models from the Sentence-BERT family. We also automatically expanded the initial dataset and fine-tuned a pre-trained Sentence-BERT model on the idiom/context matching task. This approach achieved the highest MRR score of 0.507. The data and the trained model are publicly available.