Inference from Longitudinal Data by Clustering and Machine Learning
摘要
We aim to propose a new approach to inferring from longitudinal data when their number is insufficient for building a model. The idea is to group samples of longitudinal curves and possibly covariates associated with them into clusters of similar elements. Cluster membership plays a similar role to the level of a qualitative factor in the model, i.e., it selects a more accurate version of the phenomenon description. When a new portion of data is collected, it is firstly classified into one of the clusters and processed in the context provided by historical observations already grouped in this cluster. These observations were previously used to select a classifier that provides recognitions that are close to those provided by the clustering. In other words, the classifier is trained on the labels provided by the clustering of historical data. The aim of classifying a newly observed object to one of the clusters is to make decisions (take actions) that were previously attached to objects in that cluster, e.g., selecting a kind of treatment. The proposed approach is illustrated by a nonparametric inference from vibrations of an operator’s cabin mounted in a large mechanical structure. The aim of inference is to select more appropriate dumping influence.