Handling Missing Data in Longitudinal Anthropometric Data Using Multiple Imputation Method
摘要
Diabetes mellitus, a prevalent and an ever-growing metabolic problem, is a widespread global health challenge. Type 2 diabetes is traditionally attributed to genetic factors and unhealthy lifestyle that can lead to obesity. Recent research has shown intrauterine fetal programing as an additional risk factor. To investigate fetal programing of diabetes in Indians, Pune Maternal Nutrition Study (PMNS) was set up by the Diabetes Unit of KEM Hospital, Pune, in 1993, in six villages near Pune. The objective was to investigate determinants of fetal growth and study the lifecourse evolution of phenotype of diabetes. The children born in the study and their parents have been serially followed-up for their growth, development, and cardiometabolic risk factors. A large dataset of over 5000 variables is created over 30 years, including demographics, anthropometry, socioeconomic status, nutrition intake, cardiometabolic risk factors, etc. Investigation of the dataset revealed a substantial number of missing values which would create an impedance in performing analytics. Hence, it was decided to impute the missing value and prepare the data for analysis in the first phase of the project. It was decided to focus initially on only 177 columns pertaining to anthropometry. To impute the missing values in the longitudinal dataset, popular algorithms like K-Nearest Neighbors and Multiple Imputation by Chained Equation (MICE) were applied. The imputed values generated from both the algorithm were compared and it was found that MICE excelled in maintaining the temporal coherence of the dataset. In conclusion, this imputation exercise underscores the paramount significance of preserving temporal consistency in longitudinal research, specifically when reporting long-term health outcomes.