Disambiguation of Russian Homographs with Transformers
摘要
The purpose of this work was to test BERT transformer-based models for homograph disambiguation in Russian—a long-standing issue in Text-To-Speech systems. The paper presents different types of Russian homographs and offers an in-depth analysis of existing methods for their disambiguation. A dataset of contexts from the Russian National Corpus for 28 homograph pairs was created and manually annotated. Three BERT models for the Russian language were selected and tested in two experiments. The results have shown that these models could achieve and outperform SOTA results in disambiguating homographs of all types on a relatively small training dataset. The pretrained models could also be used to disambiguate new pairs of intraparadigmatic homographs absent from the original dataset.