Integration of Remote Sensing and Meteorological Data for Rapid Sugarcane Yield Estimation Using Machine Learning
摘要
Sugarcane is a major commercial crop in India that requires careful monitoring due to its substantial contribution to the Indian economy and employs many people through the sugar and ethanol production industries. This study developed a rapid workflow for sugarcane yield estimation by integrating remote sensing and meteorological datasets through machine learning (ML). The seven ML algorithms namely, “random forest (RF),” “support vector regression,” “decision tree (DT),” “K-nearest neighbors (KNN),” “XGBoost (XGB),” “Gradient boosting regression,” “ridge regression” and one “customized stack ensemble model” were used predict the sugarcane yield under two different datasets condition: integrated MODIS (Moderate resolution imaging spectroradiometer) data with meteorological data and standalone meteorological data. Monthly mean value of eight MODIS data products, including “NDVI (Normalized difference vegetation index),” “EVI (Enhanced vegetation index),” “GPP (Gross primary productivity),” “FPAR (Fraction of photosynthetically active radiation),” “LAI (Leaf area index),” “NDWI (Normalized difference water index),” “VCI (Vegetation condition index),” and “ET (Evapotranspiration),” were acquired over a 21 years period to be used as a feature in models. The classification accuracy of the sugarcane class was measured using the F1 score, which was 88.23%. The result revealed that alone meteorological data was not sufficient for accurate estimation of yield. However, a significant improvement in estimation accuracy was observed when meteorological data was integrated with MODIS data. Based on the accuracy metric, the stack ensemble model provided more prediction accuracy (RMSE of 2.01 t/ha) followed by XGB (RMSE of 3.81 t/ha), DT (RMSE of 4.36 t/ha), and RF (RMSE of 4.27 t/ha). Based on the feature important output provided by algorithms, features of FPAR, VCI, GPP, and Z41 (morning relative humidity) were found to be significant contributors to the prediction model. The entire study workflow can be applied in various geographical regions and adapted for estimating yields in different crops as well.