Improving Anthropometric Datasets: Comparison of Machine Learning and Statistical Approaches to Fill Data Gaps
摘要
Digital Human Models (DHM) support virtual ergonomic design and evaluation, but require accurate scaling of physical dimensions based on anthropometric datasets. However, data gaps in anthropometric datasets can pose challenges and limit their usability. This study examined and compared the performance and feasibility of linear regression and several machine learning models (support vector machines (SVM), k-nearest neighbors (KNN), and gradient boosting (GB)) for approximating missing anthropometric data. Two datasets were used: the Study of Health in Pomerania (SHIP) from Germany, which represents a working population, and the ANSUR II dataset, which includes US military personnel. Model training was performed for eleven anthropometric dimensions (i.e., evaluation parameters), using the remaining anthropometric dimensions as features. To improve the comparability between the two datasets, virtual datasets were generated using a synthesis algorithm to harmonize the number of subjects in both datasets. Model performance was assessed using the root mean square error (RMSE) and compared to allowable error (AE) thresholds defined in ISO 20685-1. Results revealed that SVM and linear regression outperformed KNN and GB. However, except for one anthropometric dimension within the ANSUR II dataset, RMSE values exceeded AE thresholds, suggesting that the investigated models were not sufficient to fill data gaps. Further analysis showed that model performance was influenced by dataset characteristics, with the military dataset yielding lower errors, likely due to its relatively greater homogeneity. Future research will focus on more sophisticated modeling approaches and broader anthropometric dataset integration to enhance compliance with ISO standards and generalizability of results.