On the entropy of written Gedeo Language
摘要
To achieve truly efficient transmission and storage of information, or at the very least to determine if such efficiency is optimal, it is essential to estimate the information content of the target source. In the realm of information theory, this measure is referred to as entropy. This study computes the entropy of written Gedeo text to assess its information abundance and predictability. Entropy, a fundamental concept in information theory, quantifies the uncertainty or randomness in a dataset, and in the context of language, it reflects the efficiency of encoding and the structure of the language. A statistical study was carried out with an n-gram natural language model trained on sample Gedeo corpora. Using an n-gram natural language model, we present the first estimate of the average information content of this language, which is found to be approximately 2.8438 bits per character, with a range between [1.4606, 4.2227] bits per character. The findings help to quantitatively characterize Gedeo language, which will be useful for applications such as language modeling, text compression, and digital resource development for low-resource languages.