Explainable machine-learning workflow to improve real-time lithofacies identification using mudlogging data from a heterogeneous clastic reservoir sequence
摘要
Lithofacies identification is essential for reservoir characterization and development, and wellbore mud-logging drilling data (DD) and gas while drilling data (GWD) can provide this during wellbore construction. Explainable machine learning (ML) techniques are applied to improve real-time lithofacies identification based on DD and GWD datasets from the heterogeneous Trias Argileux Grèseux Inférieur (TAGI) sandstone reservoirs from three wellbores in the Sif Fatima oil field, Berkine basin (South Algeria). Data from two wells are used to train the ML models with the other well providing a blind test for the trained models. Shapely additive explanations (SHAP), permutation feature importance (PFI), and local interpretable model-agnostic explanations (LIME) techniques are applied to separately identify the relative influence of each DD and GWD variable on lithofacies predictions. Combinations of these explainable ML techniques have not been applied previously to predict lithofacies in heterogeneous reservoirs. The XGBoost model achieved the best lithofacies prediction accuracy (95.2% on testing data and 93.1% on validation data) outperforming a random forest model. The SHAP results provided more reliable feature influence information identifying propane (C3), bit run (BitRun), and rotation per minute (RPM) as the three input variables exerting the most influence on lithofacies predictions. Evaluating DD and GWD datasets with XGBoost combined with SHAP, PFI, and LIME techniques for feature selection provides useful insights for enhancing drilling operations in heterogeneous clastic reservoirs. The overlap in lithofacies properties in TAGI heterogeneous reservoirs leads to interpretational ambiguity, highlighting the need for such models to reduce uncertainty when data is limited.