Imbalanced multi-label datasets are one of the bottlenecks of machine learning models. The imbalanced distribution of labels in these datasets are a result of low numbers of the minority class which in most cases lead to biases in model predictions. Algorithms favor the majority class while leaving the minority class to poor generalization by the model. Attempts to balance the disparity within the classes individually creates more issues across the other classes. This paper introduces the Multiview learning approach for our previously created transformer model, SOCIALDISTILBERT which combines pretrained large language models and embeddings augmented with techniques such as MLeNN, MLSOL, and GAN. This helps address the issue of imbalanced multi-label datasets in classification. This framework combines the original tokenized text, and the augmented embeddings extracted from the penultimate layer of the transformer giving the model the ability to learn from both sources of information. This approach provides the opportunity to conserve the contextual representations of the input text. It also makes it possible for training transformers with augmented embeddings and improves the issue of imbalance multi-label datasets. This research work contributes to the growing body of research in mitigating the problem of data imbalance in multi-label datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving SOCIALDISTILBERT with Augmented Embeddings

  • Michael Abobor,
  • Darsana P. Josyula

摘要

Imbalanced multi-label datasets are one of the bottlenecks of machine learning models. The imbalanced distribution of labels in these datasets are a result of low numbers of the minority class which in most cases lead to biases in model predictions. Algorithms favor the majority class while leaving the minority class to poor generalization by the model. Attempts to balance the disparity within the classes individually creates more issues across the other classes. This paper introduces the Multiview learning approach for our previously created transformer model, SOCIALDISTILBERT which combines pretrained large language models and embeddings augmented with techniques such as MLeNN, MLSOL, and GAN. This helps address the issue of imbalanced multi-label datasets in classification. This framework combines the original tokenized text, and the augmented embeddings extracted from the penultimate layer of the transformer giving the model the ability to learn from both sources of information. This approach provides the opportunity to conserve the contextual representations of the input text. It also makes it possible for training transformers with augmented embeddings and improves the issue of imbalance multi-label datasets. This research work contributes to the growing body of research in mitigating the problem of data imbalance in multi-label datasets.