Comparison of Three Machine Learning Approaches in Determining Total Organic Carbon (TOC): A Case Study from Marcellus Shale Formation, New York State
摘要
Total Organic CarbonTotal Organic Carbon (TOC) containedTOCTotal Organic Carbon by subsurface source rock units is an ideal parameter for predicting the potential production of gas and oil shales because it primarily relates to the organic matter content. When discussing drilling operations for gas shale, generating an accurate TOC depth distribution is one of the most important steps in the process of estimating the gas abundance. TOC is measured customarily by retrieving core samples from the boreholes and analyzing them in specialized labs. Unfortunately, TOC measurements are not consistently recorded in drilling operations. Also, the TOC measurements on samples from a borehole are discrete rather than continuous. One way to solve these issues is to employ the capabilities of machine learningMachine learning (ML) approaches. We present a novel comparative study of the performance of three different ML methodologies in generating TOCTotal Organic Carbon content of the Marcellus ShaleMarcellus shale in New York State: Multilayer Perceptron Neural Networks (MLPNN)Multilayer Perceptron Neural Networks (MLPNN), Genetic AlgorithmsGenetic Algorithms (GA) (GA), and Support Vector MachinesSupport Vector Machines (SVM) (SVM). Estimating and evaluating the intelligent models of predicted TOC divides the data set of the well logs into three steps: training, validation, and application step. The input data are gamma ray, neutron porosity, and bulk density logs, with the expected output being TOC log. We built three models using MLPNN, SVM, and GA methods, with both normalized and non-normalized input data. Evaluation of statistical model performance indicates that the best method for creating the synthetic TOC log is GA using normalized data, with Regression (R) of 0.89, Normalized Mean Square Error (nMSE) of 0.35, Mean Square Error (MSE) of 1.0, and Mean Absolute Error (MAE) of 0.48. Modeling with normalized input data provided better results than modeling with non-normalized data.