Forecasting area and yield of cereal crops in India: intelligent choices among stochastic, machine learning and deep learning techniques
摘要
This study explores the dynamics and performance of different time series models, such as stochastic, machine learning and deep learning models. The need to investigate different techniques in the process of selecting the best model for forecasting area and yield of cereal crops in India and thereby yielding efficient predictions is demonstrated in this study. Six different time series models, namely, the autoregressive integrated moving average (ARIMA), artificial neural network (ANN), support vector regression (SVR), random forest (RF), convolutional neural network (CNN) and long short-term memory (LSTM) models, have been applied to predict the areas and yields of major cereal crops, such as paddies, wheat and maize. Data ranging from 1966 to 2021 concerning major states of India with respect to each cereal crop were considered for the study. Error metrics such as the root mean square error, mean absolute percentage error and mean absolute error were used to capture the best performing models. Among the 63 series analyzed, the CNN is found to be the best-performing model for 27% of the datasets, followed by LSTM (21%), SVR (16%), ARIMA (16%), RF (11%), and ANN (9%). The study concluded that deep learning models, particularly CNN and LSTM, are superior in adaptability and robustness for capturing patterns in cereal data. Furthermore, the complex nature of the dataset often favours deep learning models, as they have proven effective in handling intricate relationships and nonlinear patterns.