Evaluation of Language Models for Multilabel Classification of Biomedical Texts
摘要
The continuous increase of data availability and the need for their utilization make it imperative to organize them into categories. Recent classification problems often involve the prediction of multiple labels simultaneously applying to a single instance. In this paper, we propose a structured approach for the implementation and evaluation of multilabel classification tasks in the context of biomedical texts. This involves selecting appropriate datasets and models, designing experiments, and defining metrics that accurately measure the models’ performance across various aspects of the task. Our results yield notable scores and conclusions for the behavior of some state-of-the-art language models in specific data. It is shown that the complexity of biomedical data and the intricacy of multilabel classification require careful consideration of these models’ capabilities to handle large label spaces, label correlations, and the nuances of biomedical language.