Predicting shear-wave velocity from few well logs plus derivatives and volatilities optimized by dual-objective feature-selection machine learning
摘要
Well-log attributes incorporating derivatives and volatility calculated for just two or three recorded well logs can be used effectively to predict shear-wave velocity (Vs) in sparsely logged wellbores. Sensitivity analysis is used to configure the well-log attribute calculations to suit specific datasets. Recorded gamma ray (GR), bulk density (PB), and compressional wave velocity (Vp) data from the Bakken Formation (North Dakota, USA) type-well logs (1000 data records) are used to demonstrate the feasibility of the method. Data-matching machine-learning algorithms K-nearest neighbor (KNN) and transparent-open-box (TOB) outperform linear regression least absolute shrinkage and selection operator (LASSO) and extreme gradient boosting (XGBoost) in Vs prediction with this dataset. The comparative methods used to assess Vs prediction performance are statistical measures derived from multi-K-fold cross-validation analysis (3-, 4-, 5-, 10-, and 15-fold) that facility uncertainty analysis. Dual-objective optimized feature selection with a KNN–sine–cosine optimizer identifies the best performing recorded log and attribute combinations for Vs prediction. An eight-variable combination including GR, PB, and Vp recorded well logs with five computed attributes enabled the TOB model to predict Vs with a mean absolute error (MAE) of ~6 m/s and root mean square error (RMSE) of ~19 m/s. A 7-variable combination including PB and Vp recorded well logs with five computed attributes enabled the KNN model to predict Vs with MAE of ~8 m/s and RMSE of ~28 m/s. Data mining of these datasets with the TOB model revealed that just a few outlying predictions associated with the stratigraphic member transition zones were responsible for the relatively high RMSE values. As there are many sparsely logged wellbores, particularly in shale plays and in reservoir development wells, the well-log attribute technique described and evaluated offers potential to enhance the spatial distribution of reliable Vs values.