In the context of training sets for Machine Learning, we use the term Overlapping Asymmetric Datasets (OADs) to refer to a combination of data shapes where a large number of observations (Vertical data, \(\mathcal {V}\) ) are described using only few features (x), and a small subset of the observations (Horizontal data, \(\mathcal {H}\) ) are described by a larger number of features (x plus some new z). A common example of such a combination is a healthcare dataset where the majority of patients are described using a baseline set of clinical and socio-demographic features, and a handful of those patients have a richer characterisation, having undergone further testing . Given a classification task, a model trained solely on \(\mathcal {H}\) will benefit from the many features, but its performance will be limited by a small training set size . In this paper we study the problem of maximising model performance on \(\mathcal {H}\) , by leveraging the additional information available from \(\mathcal {V}\) . Our approach is based on the notions of stacked generalization and meta-learning, where the predictions generated by an ensemble of weak classifiers for \(\mathcal {V}\) are fed into a second-tier meta-learner, where the z features are also used. We conduct extensive experiments to explore the benefits of this approach over a range of dataset configurations. The results suggest that stacking improves model performance, while using z features only provides modest improvements. This may have practical implications as it suggests that in some settings, the effort involved in acquiring the additional z features is not always justified.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stacked Generalization for Overlapping Asymmetric Datasets

  • Matthew McTeer,
  • Paolo Missier

摘要

In the context of training sets for Machine Learning, we use the term Overlapping Asymmetric Datasets (OADs) to refer to a combination of data shapes where a large number of observations (Vertical data, \(\mathcal {V}\) ) are described using only few features (x), and a small subset of the observations (Horizontal data, \(\mathcal {H}\) ) are described by a larger number of features (x plus some new z). A common example of such a combination is a healthcare dataset where the majority of patients are described using a baseline set of clinical and socio-demographic features, and a handful of those patients have a richer characterisation, having undergone further testing . Given a classification task, a model trained solely on \(\mathcal {H}\) will benefit from the many features, but its performance will be limited by a small training set size . In this paper we study the problem of maximising model performance on \(\mathcal {H}\) , by leveraging the additional information available from \(\mathcal {V}\) . Our approach is based on the notions of stacked generalization and meta-learning, where the predictions generated by an ensemble of weak classifiers for \(\mathcal {V}\) are fed into a second-tier meta-learner, where the z features are also used. We conduct extensive experiments to explore the benefits of this approach over a range of dataset configurations. The results suggest that stacking improves model performance, while using z features only provides modest improvements. This may have practical implications as it suggests that in some settings, the effort involved in acquiring the additional z features is not always justified.