<p>Today, it is particularly important to recognize the semantic similarity between texts in different languages due to the emergence of new natural language processing models like ChatGPT and Bard. These models can provide more accurate and comprehensive answers to users’ questions by identifying semantic similarity between two texts in different languages. Cross-lingual semantic similarity refers to the process of calculating similarity between two pieces of text in different languages. This paper aims to present an improved method for finding similarities between sentences in different languages. Some of the current methods create the same vector space to achieve this, while others use machine translation to translate the text into another language and then determine similarity between the two sentences using monolingual sentence similarity methods. To provide finer semantic distinction, the degree of similarity is expressed as a number between 0 and 5. Over the past few years, the progress in language models based on transformers has paved the way for improvements in detecting text similarity. This article discusses the utilization of ensemble models with transformers to determine the semantic similarity of sentences in Persian and English languages utilizing the Persian-English corpus. According to our findings, this ensemble approach achieves a Pearson correlation of 95.28% with the reference (ground truth) similarity scores, outperforming state-of-the-art methods such as the ECNU ensemble (74.93%) and SEF@UHH (53.84%). These results indicate that our method surpasses previous techniques for discovering similarities between sentences in different languages, particularly for low-resource language pairs like Persian-English. The key innovations of our approach include (1) eliminating the need for machine translation, (2) combining multiple transformer models to create a shared vector space, and (3) using a neural network-based scoring mechanism to capture nonlinear relationships between sentence vectors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble transformer for cross-lingual semantic textual similarity

  • Mohammad Abdous,
  • Poorya Piroozfar,
  • Behrouz Minaei Bidgoli

摘要

Today, it is particularly important to recognize the semantic similarity between texts in different languages due to the emergence of new natural language processing models like ChatGPT and Bard. These models can provide more accurate and comprehensive answers to users’ questions by identifying semantic similarity between two texts in different languages. Cross-lingual semantic similarity refers to the process of calculating similarity between two pieces of text in different languages. This paper aims to present an improved method for finding similarities between sentences in different languages. Some of the current methods create the same vector space to achieve this, while others use machine translation to translate the text into another language and then determine similarity between the two sentences using monolingual sentence similarity methods. To provide finer semantic distinction, the degree of similarity is expressed as a number between 0 and 5. Over the past few years, the progress in language models based on transformers has paved the way for improvements in detecting text similarity. This article discusses the utilization of ensemble models with transformers to determine the semantic similarity of sentences in Persian and English languages utilizing the Persian-English corpus. According to our findings, this ensemble approach achieves a Pearson correlation of 95.28% with the reference (ground truth) similarity scores, outperforming state-of-the-art methods such as the ECNU ensemble (74.93%) and SEF@UHH (53.84%). These results indicate that our method surpasses previous techniques for discovering similarities between sentences in different languages, particularly for low-resource language pairs like Persian-English. The key innovations of our approach include (1) eliminating the need for machine translation, (2) combining multiple transformer models to create a shared vector space, and (3) using a neural network-based scoring mechanism to capture nonlinear relationships between sentence vectors.