Multicomponent English and Russian Terms Alignment in a Parallel Corpus Based on a SimAlign Package
摘要
Abstract
The article proposes a method for the alignment of multicomponent terminological units in scientific and technical texts placed in English-Russian parallel corpus. The approaches, methods, software, and levels of text alignment in parallel corpora are analyzed. The linguistic peculiarities of English- and Russian-language terminology influencing the process of alignment of special lexicon, as well as the peculiarities of SimAlign package operation in parallel text processing, are investigated. Two approaches to the alignment of multicomponent terms in a parallel corpus are analyzed. The influence of peculiarities of languages with different grammatical characteristics is shown.