A robust multivariate statistical framework for integrative analysis of anthropometric and biophysiological health data
摘要
High-dimensional dataset refers to datasets in which the number of variables is very large relative to the number of observations, but in this study, emphasis was on the number of response variables being larger than the number of observations. Hence, this study introduced and evaluated a new multivariate regression technique named Multivariate Joint Latent Regression (MJLR), specifically developed to tackle response variables in a high-dimensional case. Four existing techniques: Multivariate Linear Regression (MLR), Partial Least Squares Regression (PLSR), Multivariate Generalized Least Squares Regression (MGLSR), and Multivariate Adaptive Regression Splines (MARS) were discussed and used as a measure of comparison with the newly developed technique via the three model selection criteria measures: Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), and Hannan-Quinn Information Criterion (HQIC). Exploratory Data Analysis (EDA) was carried out on the real-life dataset. Descriptive statistics were employed to analyze the anthropometric variables (W1 to W17) and the bio-physiological variables (Z1 to Z83) for eighty (80) subjects. The results from the EDA using Correlation Matrix, Variance Inflation Factor (VIF), Principal Component Analysis (PCA), and Condition Index (CI) revealed that there was the presence of multicollinearity among the predictor variables, and also that the dataset is heteroscedastic. The first objective of this study was achieved by developing a new technique named MJLR, in which the predictor variables were first of all scaled to standard values, and then Singular Value Decomposition (SVD) was done on both the predictor variables and the response variables. The model was used to take into account the latent space representation of the data by removing the first few components of the SVD, and then a simplified but useful regression was done, and later RStudio was used to write a script for its full implementation. The result of the second objective, which was to evaluate the performance of the new method in comparison with the existing techniques on a dataset of anthropometric and bio-physiological variables, revealed significantly smaller values of AIC, BIC, and HQIC for the developed MJLR, as against the four existing techniques, making it more efficient. The result of the third objective, which was to analyze the simulated data of varying high dimensional sample sizes of the response variable and predictor variables for both the new and existing methods revealed that for the four different simulated studies with constant sample sizes (n = 20), number of predictor variables (p = 5), number of response variables (q) of 30–1000; (n = 40, p = 8, q = 60–1000), (n = 80, p = 16, q = 100–1000), and (n = 240, p = 48, q = 300–1000), the performance of MLR, PLSR, MGLSR, and MARS deteriorated considerably as the number of response variables was increased as seen by the increasing values of AIC, BIC and HQIC, which also is an indication that the higher the dimensionality (complexity of the data) of the response variables, the more these methods tend to over fit the data. Thus, it is only the developed MJLR whose performance was consistently appropriate and improved, even as the number of response variables increased. The result of this study contributes to advancing multivariate regression methodology by offering a highly efficient technique for response variable high-dimensional data analysis. The study thereby recommended, among others, that future research should modify MJLR or develop a new model to handle cases where both response and predictor variables are high-dimensional.