Building Multilingual Terminological Bridges Between Language-Specific Knowledge Silos
摘要
In this work-in-progress paper, we present Extractomat – our three-lingual Automated Term Extraction (ATE) framework for English, German, and Ukrainian. The framework follows a hybrid iterative approach for ATE in multilingual and cross-domain settings. The approach is tailored to extracting terms from scientific texts in scholarly domains. We report the results of our initial evaluation experiments with Extractomat over our OTRT dataset and an out-of-the-shelf Named Entity Recognition (NER) model over the English subset of the ACTER dataset. The results of these experiments are comparable to the best-performing solutions for multilingual ATE and NER. These findings indicate that the iterative combination of linguistic, statistical, and neural ATE methods, when fully integrated in Extractomat, has the potential to improve the State of the Art (SotA) in the mentioned settings.