Embedding techniques for categorical attributes play a pivotal role in the performance of machine learning inference, especially for downstream analysis of genomic sequence data. Categorical data do not convey quantitative information to directly influence the model parameters in mapping input to the desired output. Two commonly used state-of-the-art embedding techniques are one-hot embedding and binary embedding, which embed the values of the categorical attributes in vectors of 0s and 1s. One-hot treats values of a particular attribute uniformly, or in other words, it considers uniform probability distribution for the random variable of the categorical attribute. On the other hand, binary embedding assigns some ordinal characteristics to the data without considering the probability distribution. Both of these techniques do not embed the underlying semantics of the genomic features. In this study, we propose two conditional probability–based feature embedding techniques for categorical data that utilize the probability distribution of the data. Furthermore, the proposed embedded vector lengths are based on the number of classes in the training data rather than the unique values an attribute can take.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Conditional Probability-Based Feature Embedding for Genomic Sequence Data

  • Parashjyoti Borah,
  • Aparajita Dutta

摘要

Embedding techniques for categorical attributes play a pivotal role in the performance of machine learning inference, especially for downstream analysis of genomic sequence data. Categorical data do not convey quantitative information to directly influence the model parameters in mapping input to the desired output. Two commonly used state-of-the-art embedding techniques are one-hot embedding and binary embedding, which embed the values of the categorical attributes in vectors of 0s and 1s. One-hot treats values of a particular attribute uniformly, or in other words, it considers uniform probability distribution for the random variable of the categorical attribute. On the other hand, binary embedding assigns some ordinal characteristics to the data without considering the probability distribution. Both of these techniques do not embed the underlying semantics of the genomic features. In this study, we propose two conditional probability–based feature embedding techniques for categorical data that utilize the probability distribution of the data. Furthermore, the proposed embedded vector lengths are based on the number of classes in the training data rather than the unique values an attribute can take.