Automatic Amharic-Tigrigna Translation Using Statistical Machine Translation Approach
摘要
Translating between Amharic and Tigrigna, two morphologically rich and low-resource languages, presents significant challenges due to their complex morphological structures. This study considers translation using baseline and unsupervised segmentation with statistical machine translation (SMT) from Amharic into Tigrigna. The researchers utilized a dataset of 200,000 parallel sentences and 1,429,775 monolingual sentences for Tigrigna to create a language model (LM). The parallel dataset was aligned using the Giza++ toolkit, and the LM was prepared with SRILM. Morfessor was used to segment the dataset. Moses open-source SMT has been used for the experiment to train, tune, and decode. The baseline translation system scored 22.28 in the first experiment. The second experiment was done using unsupervised segmentation, which scored 25.28. These results demonstrate that morpheme-based units can substantially improve translation quality for Amharic-Tigrigna. Future research should explore supervised segmentation techniques with additional datasets to enhance translation performance.