错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation and Domain Adaptation of Similarity Models for Short Mathematical Texts

  • Christian Steinfeldt,
  • Helena Mihaljević

摘要

The ability to identify semantic similarities between short mathematical texts can reveal synergies among mathematicians with overlapping interests and enhance research infrastructure services and bibliometric databases. While various methods exist for computing textual similarity, their effectiveness in this specialized domain remains largely unexplored. Due to the lack of explicit semantic similarity datasets for mathematical texts, we formulate two classification tasks and one anchoring task based on data from the mathematical publication databases zbMATH Open and arXiv. We evaluate several models, including BERT, SBERT, and SPECTER 2, using different pooling strategies and domain adaptation techniques. Our results indicate that SBERT demonstrates superior performance in an out-of-the-box scenario, while domain adaptation significantly enhances performance for specific tasks. Furthermore, the study demonstrates the effectiveness of these text similarity models in recommending classification codes from textual data alone. These findings can be used to improve research infrastructure services, such as enhancing the search functionality of bibliometric databases, thereby aiding researchers in identifying synergies and exploring thematic cohesion within the mathematical research landscape.