Multitask classification: assessing data complexity and determining correlations with classifier performance
摘要
This paper introduces significant advancements in addressing data complexity within multitask classification problems. It presents 11 novel multitask data complexity measures adept at handling the intricacies of hybrid datasets and coping with missing values. Moreover, it introduces the Multitask Customized Naïve Associative Classifier (MCNAC), a pioneering instance-based multitask supervised classifier crafted through algorithmic adaptation. We further assess the complexity of 23 diverse datasets, revealing substantial correlations between eight of the proposed data complexity measures and the performance metrics of the novel classifier. These findings underscore the practical utility of these metrics in effectively predicting classifier performance across varied dataset scenarios. Furthermore, the MCNAC emerges as a promising solution, demonstrating its versatility in accommodating hybrid and incomplete data while retaining interpretability. In addition, we used an automatic attribute weighting strategy for MCNAC based on differential evolution. The experimental analysis shows the good performance of MCNAC in comparison with state-of-the-art multitask classifiers. In summary, this research highlights the critical significance of meticulously evaluating data complexity in multitask supervised classification tasks. By introducing a comprehensive suite of measures and innovative classifiers precisely engineered to address these complexities, our study contributes significantly to the enhancement of performance and interpretability in multitask classification contexts, thereby enriching the landscape of machine learning methodologies.