错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning and network-based integration of qualitative and quantitative characters for decoding the yield architecture of rice (Oryza sativa L.)

  • Soham Hazra,
  • Subhadwip Ghorai,
  • Suvojit Bose,
  • Avishek Chatterjee,
  • Ankur Mukhopadhyay,
  • Pabitra Kumar Ghosh,
  • Sourav Roy,
  • Rajdeep Mohanta,
  • Poulomi Sen

摘要

Rice (Oryza sativa L.) harbors substantial yet underutilized diversity for yield improvement necessitating novel strategies to utilize the diversity and enhance yield enhancement. Conventional analyses generally treat qualitative descriptors and quantitative yield components separately, thereby obscuring the holistic plant ideotype. This study integrated 143 qualitative morphological descriptors with 14 quantitative characters recorded on 62 rice genotypes which included landraces from West Bengal and Tripura and cultivated varieties to decipher the architecture of grain yield. An explicit multistage analytical pipeline was implemented. A global character association network described the high dimensional topology of interactions between qualitative and quantitative characters. Relative importance of the characters towards yield per plant was deciphered through Random Forest regression which identified the 20 most influential characters. Network reconstruction and correlation analysis of the most important characters revealed that panicle weight, 1000 seed weight and tillers per plant were the key quantitative characters. Among the qualitative characters erect flag leaf and medium panicle density were most sensitive towards yield per plant. Multivariate analysis using a three-dimensional Principal Component Analysis (3D PCA) classified the 62 genotypes into four distinct genetic clusters. A selection index derived from the top characters identified 10 elite ideotypes that combined heavy panicles with efficient plant architecture. The results demonstrated that qualitative morphotypes functioned as critical drivers of yield formation rather than merely descriptive markers. Integrating the characters into machine learning and network frameworks substantially enhanced the precision of selection in prebreeding programs targeting indigenous rice germplasm.