Efficient Data Preprocessing for Ecological Quality Assessment in Marine Environments
摘要
The integration of machine learning (ML) to predict the Biotic Index (BI) for assessing Ecological Quality (EQ) of marine environments using environmental DNA (eDNA) metabarcoding data represents a significant advance in biomonitoring. However, the complexity of these data types poses challenges like systematic variability and the curse of dimensionality, impacting prediction quality. To address these issues, we propose a generic ML pipeline with crucial preprocessing steps, including normalization and dimensionality reduction techniques, to enhance prediction quality. Our work involves comparing the performance of the Random Forest Classifier for predicting EQ classes across seven markers. We examine various combinations of two normalization techniques and a suitable dimensionality reduction with an optimal number of reduced features. Through rigorous experimentation, we identify the most effective combination and optimal number of components to retain, establishing a standardized preprocessing protocol for this data type. This robust framework for preprocessing metabarcoding data significantly advances EQ prediction.