<p>Identifying essential proteins is vital for understanding cellular functions, discovering new drug targets, and linking genes to human diseases. Traditional methods like gene deletions and RNA interference are often slow and resource-intensive. To address this, computational approaches analyze sequence-based features and topology-based features from protein–protein interaction (PPI) networks. However, these methods face challenges due to imbalances in PPI networks, leading to less accurate predictions. This study aims to improve the identification of essential proteins by combining gene ontology (GO) annotations for yeast proteins with topological information from PPI data and applying deep learning techniques. In the feature extraction process, the PPI network is converted into a graph, and node embeddings are generated using the Node2Vec algorithm, which captures structural and relational properties within the network. Additionally, 695 GO features were collected from the UniProt database and merged with the node embeddings to form a comprehensive feature set. This integrated approach, termed Node Embedding + GO features, significantly enhances prediction accuracy while addressing data imbalance using SMOTE. The proposed method achieved an accuracy of 94.20%, providing a robust and scalable framework for future research in essential protein identification. The study of essential proteins is crucial for understanding cellular functions, identifying potential drug targets, and advancing disease research. It has significant practical applications in drug development.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning for predicting essential proteins using topological features and gene ontology

  • Md. Shahidul Islam,
  • Md. Rafiqul Islam,
  • Durjoy Mistry

摘要

Identifying essential proteins is vital for understanding cellular functions, discovering new drug targets, and linking genes to human diseases. Traditional methods like gene deletions and RNA interference are often slow and resource-intensive. To address this, computational approaches analyze sequence-based features and topology-based features from protein–protein interaction (PPI) networks. However, these methods face challenges due to imbalances in PPI networks, leading to less accurate predictions. This study aims to improve the identification of essential proteins by combining gene ontology (GO) annotations for yeast proteins with topological information from PPI data and applying deep learning techniques. In the feature extraction process, the PPI network is converted into a graph, and node embeddings are generated using the Node2Vec algorithm, which captures structural and relational properties within the network. Additionally, 695 GO features were collected from the UniProt database and merged with the node embeddings to form a comprehensive feature set. This integrated approach, termed Node Embedding + GO features, significantly enhances prediction accuracy while addressing data imbalance using SMOTE. The proposed method achieved an accuracy of 94.20%, providing a robust and scalable framework for future research in essential protein identification. The study of essential proteins is crucial for understanding cellular functions, identifying potential drug targets, and advancing disease research. It has significant practical applications in drug development.