Shale Lithofacies Identification Based on Machine Learning Classification Algorithm: A Case Study of the Lianggaoshan Formation in the Jurassic Period, Sichuan Basin
摘要
Determining the lithofacies of the target interval is important work to help evaluate shale reservoirs, explore the spatial distribution of shale oil, and optimize exploration targets. Conventional logging lithofacies identification mainly relies on parametric equations, crossplots, or cluster analysis. The identification speed is fast, but the accuracy is relatively low. In order to enhance the effectiveness and precision of lithofacies identification, four commonly used machine learning algorithms are employed to construct a shale lithofacies classification model. Utilizing total organic carbon content, mineral composition, and bedding structure as criteria, the shale within the Lianggaoshan Formation in the Sichuan Basin is categorized into eight distinct lithofacies. Parameters such as gamma, density, resistivity, and radioactive elements (uranium and thorium) in logging are selected to form a sample dataset. After data preprocessing and standardization, support vector machines (SVM), random forests (RF), multi-layer perceptrons (MLP), and gradient boosting decision trees (XGBoost) supervised machine learning algorithms are used to compare their effects. The findings from the experiments indicate that the classification accuracy of the shale lithofacies using the particle swarm optimization multi-layer perceptron algorithm for the Lianggaoshan Formation shale (82%) surpasses that of the support vector machine (75%), random forest (74%), and gradient boosting decision tree (78%). The perceptron classification model has high discrimination and accurate recognition in single-well lithofacies recognition. It has obvious advantages in multi-classification problems compared to other models. This may be related to the imbalance of different algorithms when facing datasets and their ability to learn a few categories.