AI is becoming state-of-the-art across scientific fields, giving novel solutions to age-old problems. In genomic prediction, Machine Learning methods could not outperform linear regressions in a general way yet, but are becoming closer. An important feature when working with genomic data, which is non other than a long sequence of information, is to account for the linkage disequilibrium, i.e. dependencies between genome variations that do not need to be close in the genome, and variate with respect to the reference genome. To explode this feature, we evaluate a Transformer trained in a small yeast dataset. Although it did not outperform the state-of-the-art results yet, the model got close achieving an \(R^2\) score of 0.389 and 0.400 in Lactate and Lactose ambients, respectively, comparing to the \(R^2\) score of 0.568 and 0.582 for Lactate and Lactose ambients, for the linear model of Lasso, proposed by [7]. This proves that there is still room for improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformers for Genomic Prediction

  • María Inés Fariello,
  • Graciana Castro,
  • Romina Hoffman,
  • Mateo Musitelli,
  • Diego Belzarena,
  • Federico Lecumberry

摘要

AI is becoming state-of-the-art across scientific fields, giving novel solutions to age-old problems. In genomic prediction, Machine Learning methods could not outperform linear regressions in a general way yet, but are becoming closer. An important feature when working with genomic data, which is non other than a long sequence of information, is to account for the linkage disequilibrium, i.e. dependencies between genome variations that do not need to be close in the genome, and variate with respect to the reference genome. To explode this feature, we evaluate a Transformer trained in a small yeast dataset. Although it did not outperform the state-of-the-art results yet, the model got close achieving an \(R^2\) score of 0.389 and 0.400 in Lactate and Lactose ambients, respectively, comparing to the \(R^2\) score of 0.568 and 0.582 for Lactate and Lactose ambients, for the linear model of Lasso, proposed by [7]. This proves that there is still room for improvement.