Towards Climate-Smart Agriculture: A Transferable Big Data and Clustering Framework for Crop Risk and Yield Prediction Using Indian Agricultural Data
摘要
Agriculture in South Asia is highly vulnerable to climate variability, soil degradation, and resource inefficiencies. To address these challenges, this study develops a big-data framework that integrates unsupervised clustering and supervised machine-learning models for crop yield and risk prediction, using Indian agricultural data spanning 15 years (2000–2014). The dataset comprises 246,092 records across 29 states and 600 + districts, including crop types, cultivated area, production, soil nutrients (N, P, K), soil type, and climatic variables. Clustering methods (K-Means, Hierarchical, DBSCAN) identified crop-soil-climate groupings, while predictive models (Ridge, Lasso, Random Forest, Gradient Boosting) were used to estimate yields. Specifically, four distinct agro-climatic zones were identified, representing unique crop-soil-climate groupings derived from the integrated dataset. Results show Random Forest achieved the best performance (R2 up to 0.88). Cluster-specific models further improved accuracy, highlighting strong interactions between soil, climate, and crop features. The framework provides decision-support insights for farmers (crop prioritization), insurers (risk-based premiums), and policymakers (fertilizer subsidy targeting). These insights aim to optimize resource allocation, guiding stakeholders in selecting crops, designing insurance premiums, and targeting subsidies efficiently. Limitations include a restricted data scope due to missing irrigation and pest incidence data, as well as the exclusion of post-2014 climate variability; future work will integrate satellite and IoT data for real-time climate-smart agriculture.