Fitting Big Data Using Structural Measurement Error Model
摘要
This article aims to introduce Structural Measurement Error Models (SMEM) in the context of big data analytics. The maximum likelihood estimation method was implemented to estimate unknown parameters and fit data to the suggested models. The performance of fitting big data using a sub-sampling algorithm was evaluated based on several criteria including the program execution time, bias, mean squared error of estimators, Akaike Information Criterion (AIC), and Bayesian Information Criterion (BIC). R software was used to conduct Monte Carlo simulation experiments. Then the research idea was implemented on real data to analyze the relationship between the number of passengers and population size, based on a sample of one million passengers. Both simulation experiments and real data analysis revealed that dividing data into several groups improved machine performance in terms of time but with less informative results in terms of AIC and BIC.