Haplotype-based autoencoders can reduce the dataset dimension and estimate haplotype block effects in different crop species
摘要
In plant breeding, many studies currently investigate the application of machine learning (ML) to genomic prediction, hoping for an improvement in prediction accuracy compared to standard models like Genomic Best Linear Unbiased Prediction (GBLUP). However, ML algorithms require much higher computational resources. This study aims to reduce the computational requirements and speed up training time by developing a novel autoencoder architecture inspired by haplotype blocks. Our approach incorporates prior knowledge on genetic linkage, inspired by haplotype block building, into the autoencoder architecture, resulting in a new encoded variable per haplotype block. We further modified our model into a semi-supervised version by adding available yield information. We used features extracted from the autoencoder’s block layer as inputs for Random Forest and GBLUP models to predict the yield of hybrid and inbred crops.
ResultsGenomic prediction based on the extracted features maintained prediction accuracies equal to using the original marker data, even with a variable reduction of up to 98% and significantly reduced computation time. Prediction accuracy of the supervised component was in some cases equal to and in some lower than the prediction accuracy achieved using GBLUP. Effects estimated for haplotype block variants using our new method showed a high correlation to the blockwise sum of marker effects, which is the current standard approach for haplotype block effects. Correlation between the two block effect estimation approaches was very low for some blocks, which might indicate the incorporation of non-linear effects by the autoencoder.
ConclusionsOur approach introduces a new perspective on processing haplotype blocks for genomic prediction, potentially providing more flexible modelling opportunities without the use of multiple binary dummy variables for each block variant. Additionally, training time for ML models may be significantly reduced by using the reduced feature sets generated using our method. By adding the semi-supervised component, the model is able to estimate values similar to marker effects for each block on yield. In future work, this may provide a new way of quantifying the importance of haplotype blocks for selection and breeding.