Streamlining Protein Data Pre-processing Assisted by Machine Learning
摘要
Proteins are the building blocks of life that play a crucial role in every biological process, serving as enzymes, structural components, signaling molecules and regulators within living organisms. Proteins are studied in regard to their structures and sequences stored int the corresponding datasets. These datasets often encounter challenges which significantly impact downstream analysis. Despite advancements in protein structure research, data pre-processing aspects remain unexplored. In this study, we address these gaps by proposing a comprehensive framework to handle common issues in protein datasets. We introduce robust data pre-processing techniques to manage missing values, categorical features and label encoding, ensuring data integrity and consistency. These techniques aim to streamline the pre-processing phase, facilitating more accurate and reliable analysis of protein dataset by employing machine learning (ML) techniques. The proposed methods are evaluated to show improvement in overall quality of data-driven studies in protein structure research, paving the way for more robust and insightful prediction applications. The evaluation of the proposed model shows that the model achieves above 90% of area under the curve (AUC) which indicates discriminate ability for classification.