Machine Learning Based Recognition Model for Lithology Classification
摘要
To enhance the accuracy and intelligence of lithology identification under complex geological conditions, this study proposes a well-logging lithology classification model based on machine learning, addressing issues such as imbalanced lithological distribution and high-dimensional logging features. The original logging data were first subjected to cleaning and standardization. Multiple well-logging parameters, including gamma ray, acoustic slowness, and neutron porosity, were selected as feature variables, while lithology types served as the target labels. The SMOTE algorithm was applied to augment the dataset and mitigate class imbalance. Correlation analysis was then employed to construct multiple diverse feature subsets. Subsequently, ten representative machine learning classification models were trained and tested. Model performance was comprehensively evaluated using four metrics: accuracy, precision, recall, and F1-score, leading to the selection of the optimal model. Furthermore, the SHAP algorithm was introduced to conduct interpretability analysis, elucidating the influence mechanisms of various logging parameters on lithology classification. Experimental results demonstrate that ensemble-based models exhibit stable performance across all evaluation metrics and outperform conventional models. The outcomes of interpretability analysis are consistent with geological understanding, confirming the model’s effectiveness and practical utility. This study innovatively integrates SMOTE-based data augmentation, multi-feature subset modeling, and model interpretability techniques, offering a high-accuracy, high-reliability approach to intelligent lithology identification from well logs. The proposed method holds significant theoretical value and engineering application potential.