Pre-Trained Language Model for Missing Value Imputation in Ocean Buoy Data
摘要
Due to factors such as equipment malfunctions and environmental disturbances, a significant amount of missing data exists in ocean in-situ observation datasets, which considerably hinders the usability of the data, especially when analyzing and predicting using the observational data. To obtain complete ocean spatiotemporal data, this study proposes an innovative pre-trained language model for ocean spatiotemporal data imputation (OSTI-PLM), which aims to leverage pre-trained language models (PLM) and graph structure to impute missing values in ocean multiple buoys in-situ observation datasets. First, we designed a spatiotemporal feature extraction layer to capture the temporal periodicity and spatial correlations of ocean data. Subsequently, a spatiotemporal tokenization module was developed to convert the spatiotemporal data into tokens, aligning with the processing requirements of the PLM. Additionally, this study introduces the Partial Freezing Attention (PFA) strategy and LoRA technique to fine-tune the PLM, further enhancing the model’s understanding of spatiotemporal data and improving imputation accuracy. Experimental results on real-world in-situ data from 20 ocean buoys in the Mediterranean demonstrate that OSTI-PLM exhibits outstanding performance in the spatiotemporal data imputation task for ocean multiple buoys datasets.