Understanding protein-protein interactions (PPIs) helps identify protein functions and develops other important applications such as drug preparation and protein-disease relationship identification. Machine learning-based approaches are being deeply researched for PPI determination to reduce the cost and time of previous testing methods. In this work, we introduce a novel ensemble learning model to predict PPIs from only protein sequences. First, the protein sequence data is converted into raw features. Then, these raw features are segmented into equal subsets before being extracted by multiple deep neural networks into local features. Concurrently, Autoencoders are used to convert the raw features into global features. Finally, these local and global features are concatenated and fed to multi-extreme gradient boosting machines to predict the interaction probability. We evaluated the proposed model using the new PPI dataset derived from the STRING database and the benchmark Human dataset. Data and source code are provided via https://gitlab.com/nhanth/EnsMF-PPI.git .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining Ensemble Learning and Multi–view Feature Extraction for Protein–protein Interaction Prediction

  • Tran Hoai-Nhan,
  • Nguyen-Phuc-Xuan Quynh,
  • Vo-Ho Thu-Sang,
  • Nguyen-Thi Lan-Anh

摘要

Understanding protein-protein interactions (PPIs) helps identify protein functions and develops other important applications such as drug preparation and protein-disease relationship identification. Machine learning-based approaches are being deeply researched for PPI determination to reduce the cost and time of previous testing methods. In this work, we introduce a novel ensemble learning model to predict PPIs from only protein sequences. First, the protein sequence data is converted into raw features. Then, these raw features are segmented into equal subsets before being extracted by multiple deep neural networks into local features. Concurrently, Autoencoders are used to convert the raw features into global features. Finally, these local and global features are concatenated and fed to multi-extreme gradient boosting machines to predict the interaction probability. We evaluated the proposed model using the new PPI dataset derived from the STRING database and the benchmark Human dataset. Data and source code are provided via https://gitlab.com/nhanth/EnsMF-PPI.git .