The Genomic Estimated Breeding Value (GEBV) is a key metric in breeding programs, guiding both selection and mating decisions. Selecting candidates based on GEBV can considerably shorten breeding cycles and enhance overall efficiency. We propose DG-Bi-LSTM, a variant encoding architecture based on Bi-LSTM, and further develop the ProxiGen-BiMAE model for GEBV prediction by incorporating Transformer mechanisms. The model is trained in two stages: first, it applies masking and unsupervised pretraining to capture the feature distribution of genomic data; second, it performs supervised training using both genomic and phenotypic information to predict GEBV. This two-stage strategy effectively addresses the challenge of high-dimensional genomic data with limited sample sizes. Experiments on Duroc pig and Colored-feather broiler datasets show that our method outperforms traditional GBLUP and BayesA by 14.92% and 11.46% in accuracy, respectively, with the pretraining phase contributing an average gain of 9.5%. Furthermore, when applied to low-cost, low-depth sequencing data, the model maintains high performance with only a 2.7% accuracy drop, indicating its potential to reduce sequencing costs while improving prediction accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Masked Bi-LSTM with Unsupervised Encoding for Genomic Breeding Value Estimation

  • Guoyu Yan,
  • Ying Ji

摘要

The Genomic Estimated Breeding Value (GEBV) is a key metric in breeding programs, guiding both selection and mating decisions. Selecting candidates based on GEBV can considerably shorten breeding cycles and enhance overall efficiency. We propose DG-Bi-LSTM, a variant encoding architecture based on Bi-LSTM, and further develop the ProxiGen-BiMAE model for GEBV prediction by incorporating Transformer mechanisms. The model is trained in two stages: first, it applies masking and unsupervised pretraining to capture the feature distribution of genomic data; second, it performs supervised training using both genomic and phenotypic information to predict GEBV. This two-stage strategy effectively addresses the challenge of high-dimensional genomic data with limited sample sizes. Experiments on Duroc pig and Colored-feather broiler datasets show that our method outperforms traditional GBLUP and BayesA by 14.92% and 11.46% in accuracy, respectively, with the pretraining phase contributing an average gain of 9.5%. Furthermore, when applied to low-cost, low-depth sequencing data, the model maintains high performance with only a 2.7% accuracy drop, indicating its potential to reduce sequencing costs while improving prediction accuracy.