错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Elevation of Enzyme Classification by Computational Intelligence: A Novel Approach of Enzymes Classification

  • Prabha Singh,
  • Sudhakar Tripathi,
  • Anand Bihari

摘要

Enzymes are complex molecules that work as catalysts in living organisms. The primary function is to accelerate the chemical reactions. Without enzymes, the different chemical reactions in a living organism would proceed very slowly. Enzymes can be classified into six classes based on their catalyzing capability of chemical reactions. Classification of enzymes is challenging due to the broad diversity of their functions and the limited structural information available for many enzymes. It is difficult to establish a clear distinction between various enzyme classes and subclasses because of complex interactions between the components of an enzyme; the categorization of a newly discovered enzyme is entirely changing. Different traditional experimental methods, such as enzyme assays and structural analysis, are often time-consuming, labor-intensive, and expensive for enzyme class prediction. On the other hand, various computational techniques like machine learning and deep learning offer a fast and cost-effective alternative. With the help of a massive dataset of enzyme sequences and structures and these computational techniques, the researchers can predict the classification and functionalities of enzymes with higher accuracy. These computational approaches are fast, more accurate, and cost-effective alternatives for classification purposes. This paper presents the efficiency of machine learning and deep learning approaches for enzyme family classification and their potential to discover their functionalities rapidly. We have employed a comprehensive dataset of protein sequences obtained from the Kaggle Data Repository. There are three classifiers: Bayes’ classifier, boosted decision tree, and decision forest model, which classify proteins in the enzyme into two classes: EC1 and EC2. For classification purposes, protein sequence information is used as an input variable. Among all these methods, boosted decision tree classification gives an accuracy of 99.2% and a recall of 99.6%. For the Bayes classifier, accuracy is 99.3% and recall of 99.8%. For the decision tree, the overall accuracy is 99.4% and the recall value is 1.