Optimizing management zone delineation technique for high-dimensional and large-volume datasets in precision agriculture
摘要
Management zone (MZ) delineation is critical for precision soil and crop management, facilitating division of fields into sub-areas with similar characteristics. This study investigated the integration of variable selection methods and clustering algorithms to improve MZ delineation accuracy using high-dimensional and large-volume datasets.
MethodsFour variable selection methods (All-Attribute [All], principal component analysis [PCA], MULTISPATI-PCA [SPCA], and t-distributed stochastic neighbour embedding [t-SNE]) and five clustering algorithms (K-means, fuzzy C-means [FCM], BIRCH, mini-batch K-means [MBK], and spatially constrained hierarchical clustering) were evaluated. Soil data encompassing eight properties (pH, organic carbon, magnesium, phosphorous, potassium, calcium, sodium, and moisture content) were collected with an on-line visible and near-infrared spectroscopy platform providing thousands of sampling points across two commercial fields. Zoning results for three, four, and five MZs were evaluated using the Silhouette and Davies-Bould inindexes. Further statistical analysis of soil properties within and between zones was performed.
ResultsThe optimal number of MZs was determined based on first-level value (FLV) frequencies, with four zones identified for Field 1 (Frequencies: 8, 14, and 8 for 3, 4, and 5 MZs) and three zones for Field 2 (Frequencies: 19, 10, and 1 for 3, 4, and 5 MZs). Further validation confirmed that t-SNE combined with FCM or K-means was most effective for Field 1, whereas t-SNE with MBK or K-means performed best for Field 2.
ConclusionIntegrating advanced variable selection and clustering, especially t-SNE with FCM, MBK, or K-means, significantly improves MZ delineation in large soil datasets. These findings provide a robust foundation for precision agriculture applications.