The optimal spatial averaging method for random/non-random missing data via super learner and its application to Tara Oceans data
摘要
Spatial averaging has become an important tool in spatial data analysis, achieved by assigning different weights to various spatial observations. Missing data is common in many spatial datasets. In this article, we propose a form of spatial average statistics for both random and non-random missing data and present the optimal sampling weights by minimizing bias and variance, respectively. The optimal weights depend on the missing mechanism. Compared to traditional logistic regression models, we consider machine learning techniques, such as Classification and Regression Trees and Random Forests, as alternative candidates for modeling the missing mechanism. We employ two simulation designs in this study: a general data generation mechanism and a simulation that mimics real data by maintaining the structure and confounding factors consistent with the real dataset. The results show that the optimal averaging strategy proposed in this paper can closely approximate the true value, regardless of the missing mechanism. Finally, we apply the proposed method to the Tara Oceans data and estimate the optimal weights for the spatial average of the relative abundance of Ammonia-Oxidizing Archaea.