Maximum Entropy Model of Synonym Selection in Post-editing Machine Translation into Kazakh Language
摘要
The work presents a model, algorithm, and experimental studies for selecting synonyms of incorrectly translated words from English into Kazakh within the framework of machine translation post-editing technology. As a model for choosing synonyms, a maximum entropy model has been developed, the distinctive feature of which is the consideration of contextual words located at any distance from the translated word in a sentence (non-consecutive collocations), which takes into account the peculiarities of the Kazakh language. A feature of the proposed solution algorithm for this model is the use of the semantic cube model proposed by the authors. The developed maximum entropy model was learned on the 250 0000 sentences parallel Kazakh-English corpus, and a test set containing 25,000 sentences was conducted. Experiments on post-editing of machine translation of the Kazakh language compared with machine translation of Google Translate showed an improvement in the BLEU metric by 6 positions.