<p>To achieve truly efficient transmission and storage of information, or at the very least to determine if such efficiency is optimal, it is essential to estimate the information content of the target source. In the realm of information theory, this measure is referred to as entropy. This study computes the entropy of written Gedeo text to assess its information abundance and predictability. Entropy, a fundamental concept in information theory, quantifies the uncertainty or randomness in a dataset, and in the context of language, it reflects the efficiency of encoding and the structure of the language. A statistical study was carried out with an n-gram natural language model trained on sample Gedeo corpora. Using an n-gram natural language model, we present the first estimate of the average information content of this language, which is found to be approximately 2.8438 bits per character, with a range between [1.4606, 4.2227] bits per character. The findings help to quantitatively characterize Gedeo language, which will be useful for applications such as language modeling, text compression, and digital resource development for low-resource languages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the entropy of written Gedeo Language

  • Eyob Tadesse Tekish,
  • Beimnet Fikadu,
  • Zenebe Shote

摘要

To achieve truly efficient transmission and storage of information, or at the very least to determine if such efficiency is optimal, it is essential to estimate the information content of the target source. In the realm of information theory, this measure is referred to as entropy. This study computes the entropy of written Gedeo text to assess its information abundance and predictability. Entropy, a fundamental concept in information theory, quantifies the uncertainty or randomness in a dataset, and in the context of language, it reflects the efficiency of encoding and the structure of the language. A statistical study was carried out with an n-gram natural language model trained on sample Gedeo corpora. Using an n-gram natural language model, we present the first estimate of the average information content of this language, which is found to be approximately 2.8438 bits per character, with a range between [1.4606, 4.2227] bits per character. The findings help to quantitatively characterize Gedeo language, which will be useful for applications such as language modeling, text compression, and digital resource development for low-resource languages.