Ensemble Machine Learning Models for Rice and Wheat Yield Prediction: A Comparative Study Across Districts in India’s Kharif and Rabi Seasons
摘要
Agriculture has been the backbone of India’s economy for centuries, providing livelihood for millions of people and stimulating economic growth. Yield assessment is essential to agricultural planning and management, including crop insurance for farmers. Traditional methods of crop yield assessment adopted in the Republic of India are based on conducting a significant number of crop cutting experiments (CCE), which are time-consuming, labor-intensive, and expensive. Regression models based on meteorological data, Earth remote sensing data, and machine learning (ML) methods can significantly reduce the number of required CCEs. This article describes estimating yields using regression models based on ML methods and ensembles. ML models based on the following methods were used: Deep Neural Networks (DNN), Random Forest, Extremely Randomized Trees, Catboost, K-Nearest Neighbors, and Linear Regression. Separate models and ensembles were built for each “season–district–crop” combination. Yield evaluation was conducted in 44 districts of Andhra Pradesh, Haryana, Jharkhand, Madhya Pradesh, Odisha, Tamil Nadu, Uttar Pradesh, Bihar, and Rajasthan. The best coefficient of determination reached 0.94. The minimum mean absolute percentage error was 6%. For most of the districts studied, the ensembling resulted in an increase of the