<p>Protein function categorization or prediction is considered to be one of the key subfields of bioinformatics and computational biology. Clarifying the relationships and functions of the proteins that essentially comprise the body is beneficial. Proteins, one of the biggest types of macromolecules, are essential to every biological function. They consist of amino acids, which are believed to be the main components of proteins. Every protein has a distinct amino acid sequence that dictates its structure and function since all proteins are made of amino acids, just as all nucleic acids are made of nucleotides. There are several ways to classify proteins, such as grouping them into several groups according to their sequences, physiological structures, and functions. Recent advances in machine learning and deep learning models have made it possible to classify proteins in a number of ways. In light of this, the primary objective of the current work was to assess how well various ML and DL algorithms performed on the Protein family database (PFAM), which serves as a reference dataset for the protein classification problem. Numerous sequences connect each protein to one of the numerous predetermined groups. Consequently, the study's findings demonstrated significant performance disparities between deep learning models like 1D Dilated CNN and machine learning models like SVM, Random Forest, and Naive Bayes. The best tested architecture was the 1D Dilated CNN, which achieved the maximum accuracy of 97%. It is among the most scalable models available, needing the least amount of time and 6%. As demonstrated in the instance of protein classification, deep learning is actually best applied in scenarios where it has the highest chance of success, such as in datasets with a high dimensionality and a large number of features.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A review on amino acid based protein classification using supervised artificial intelligence (AI) models

  • Rida Zulfiqar,
  • Talha Ahmed Khan,
  • Abdullah Ayub Khan,
  • Rubaika Akhtar,
  • Sajid Ullah

摘要

Protein function categorization or prediction is considered to be one of the key subfields of bioinformatics and computational biology. Clarifying the relationships and functions of the proteins that essentially comprise the body is beneficial. Proteins, one of the biggest types of macromolecules, are essential to every biological function. They consist of amino acids, which are believed to be the main components of proteins. Every protein has a distinct amino acid sequence that dictates its structure and function since all proteins are made of amino acids, just as all nucleic acids are made of nucleotides. There are several ways to classify proteins, such as grouping them into several groups according to their sequences, physiological structures, and functions. Recent advances in machine learning and deep learning models have made it possible to classify proteins in a number of ways. In light of this, the primary objective of the current work was to assess how well various ML and DL algorithms performed on the Protein family database (PFAM), which serves as a reference dataset for the protein classification problem. Numerous sequences connect each protein to one of the numerous predetermined groups. Consequently, the study's findings demonstrated significant performance disparities between deep learning models like 1D Dilated CNN and machine learning models like SVM, Random Forest, and Naive Bayes. The best tested architecture was the 1D Dilated CNN, which achieved the maximum accuracy of 97%. It is among the most scalable models available, needing the least amount of time and 6%. As demonstrated in the instance of protein classification, deep learning is actually best applied in scenarios where it has the highest chance of success, such as in datasets with a high dimensionality and a large number of features.