<p>Lysine 2-hydroxyisobutyrylation (Khib), a Post-Translational Modification&#xa0;(PTM),&#xa0;plays a pivotal role in regulating protein structure and function, with emerging evidence highlighting its significance in cellular metabolism, transcriptional regulation, and disease pathways. However, the experimental identification of Khib sites is hindered by labour-intensive methods and the dynamic nature of these modifications. To address this challenge, we propose a computational framework utilizing a Light Gradient Boosting Machine (LightGBM) for predicting Khib sites.&#xa0;37-amino acid peptide sequences are represented using a hybrid feature set that combines Evolutionary Scale Modeling (ESM), Composition, Transition, Distribution (CTD), and AAindex descriptors. These features capture both the evolutionary and physicochemical properties of the protein sequences. Mutual information-based feature selection enhances model performance, while LightGBM outperforms alternative classifiers, including Support Vector Machines (SVM) and&#xa0;XGBoost. Validation on <i>Homo sapiens</i>, <i>Toxoplasma gondii</i>, and <i>Oryza sativa</i> datasets yielded Area under ROC Curve (AUC) values of 0.846, 0.836, and 0.788, respectively, surpassing existing predictors such as iLys-Khib and&#xa0;KhibPred. Additionally, sequence analysis revealed species-specific amino acid preferences surrounding Khib sites, providing insights into the biological determinants of this modification and advancing the prediction of Khib sites across species.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing lysine 2-hydroxyisobutyrylation site prediction using LightGBM and hybrid sequence features

  • Heba M. Elreify,
  • Fathi E. Abd El-Samie,
  • Moawad I. Dessouky,
  • Hanaa Torkey,
  • Said E. El-Khamy,
  • Wafaa A. Shalaby

摘要

Lysine 2-hydroxyisobutyrylation (Khib), a Post-Translational Modification (PTM), plays a pivotal role in regulating protein structure and function, with emerging evidence highlighting its significance in cellular metabolism, transcriptional regulation, and disease pathways. However, the experimental identification of Khib sites is hindered by labour-intensive methods and the dynamic nature of these modifications. To address this challenge, we propose a computational framework utilizing a Light Gradient Boosting Machine (LightGBM) for predicting Khib sites. 37-amino acid peptide sequences are represented using a hybrid feature set that combines Evolutionary Scale Modeling (ESM), Composition, Transition, Distribution (CTD), and AAindex descriptors. These features capture both the evolutionary and physicochemical properties of the protein sequences. Mutual information-based feature selection enhances model performance, while LightGBM outperforms alternative classifiers, including Support Vector Machines (SVM) and XGBoost. Validation on Homo sapiens, Toxoplasma gondii, and Oryza sativa datasets yielded Area under ROC Curve (AUC) values of 0.846, 0.836, and 0.788, respectively, surpassing existing predictors such as iLys-Khib and KhibPred. Additionally, sequence analysis revealed species-specific amino acid preferences surrounding Khib sites, providing insights into the biological determinants of this modification and advancing the prediction of Khib sites across species.