Data-Driven Machine Learning Approach for Estimation of the Total Organic Carbon (TOC): Implications of Data Splitting and Model Interpretability
摘要
This study investigates the impact of data-splitting strategies on machine learning (ML)-based estimation of total organic carbon (TOC) from well log data, a critical parameter in petroleum source rock evaluation. Three ML models—a polynomial-based group method of data handling (GMDH), an ensemble random forest (RF), and a back-propagation neural network (BPNN)—were applied to predict TOC using well logs from the Mandawa Basin, Tanzania. Results indicated that data partitioning significantly influences model performance, with a 70:30 training/testing split yielding optimal results (GMDH: R2 = 0.959, RMSE = 0.065, and MAE = 0.048). Model interpretability using SHAP (SHapley Additive exPlanations) demonstrated that neutron porosity (NPHI) and sonic travel time (DT) were the most influential predictors, aligning with geological expectations. The limitations of this study, including dataset size, log quality, and model generalization, are also discussed, providing a robust framework for TOC prediction in data-limited settings.