A novel diversity-based selective ensemble method for small sample expert credibility assessment
摘要
Machine learning methods can often be used effectively to predict the credibility of experts in science and technology evaluation. However, there may be data scarcity and poor prediction performance. To address these issues, a diversity-based selective ensemble method is proposed. In this methodology, data preprocessing is first used to perform feature engineering. Then, the Gaussian mixture model (GMM) is utilized to generate virtual samples to solve small sample issues (i.e., data augmentation technique for small samples), complemented by a diversity-driven mechanism for sample filtering. Additionally, 14 diverse statistical, artificial intelligence, and ensemble models are base models. Finally, based on the hierarchical clustering algorithm, a novel selective ensemble model was proposed to improve the model’s generalization ability by fusing the model bias, variance, diversity, and complexity mechanism. A real-world expert credibility dataset was used to validate the effectiveness and feasibility of the proposed method. The experimental results demonstrated that the proposed diversity-based selective ensemble model outperforms all other models considered in this study. Moreover, sample diversity (e.g., empirical formulas and interval sampling), model diversity, parameter diversity, and data augmentation mechanisms are further analysed to verify their importance in a selective ensemble. This can be considered a promising solution for small sample expert credibility assessment.