Exploring the nuanced realm of emotions in textual content, this research conducts a comprehensive evaluation of text classification models, specifically focusing on BERT and RoBERTa as prominent language models. The dataset, representing 12 distinct emotions numerically, undergoes meticulous preprocessing and augmentation with advanced techniques from the Gensim library and FastText Embedding. Feature extraction algorithms, particularly BERT and RoBERTa, are scrutinized for their effectiveness in capturing nuanced text representations. Introducing a novel ensemble method, the research demonstrates its consistent superiority over individual models, achieving notable accuracy scores of 92.5% and 94.7% with BERT and RoBERTa, respectively. In an added layer of sophistication, hybrid feature selection involving LASSO and Recursive Feature Elimination (RFE) is incorporated post-feature extraction to refine and enhance the model’s performance. Through comprehensive comparative analyses, both pre and post hyperparameter tuning, the ensemble method asserts dominance over traditional machine learning classifiers, even surpassing Random Forest ensembles in accuracy and recall. Deep learning classifiers further validate the ensemble’s efficacy, achieving an impressive accuracy of 95.6% with RoBERTa. The study also reveals RoBERTa’s consistent slight advantage over BERT, aligning with established trends in natural language processing tasks. Beyond the research realm, these findings bear practical implications, as the proposed ensemble method, integrated with RoBERTa and enhanced through hybrid feature selection, emerges as a versatile and robust solution for accurate and nuanced emotion classification in diverse real-world scenarios, offering valuable insights for enhancing text classification performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-Based Emotion Recognition with Hybrid Feature Selection and Ensemble Classification: A BERT and RoBERTa Approach

  • Nevin Sugu,
  • Nirmal Varghese Babu

摘要

Exploring the nuanced realm of emotions in textual content, this research conducts a comprehensive evaluation of text classification models, specifically focusing on BERT and RoBERTa as prominent language models. The dataset, representing 12 distinct emotions numerically, undergoes meticulous preprocessing and augmentation with advanced techniques from the Gensim library and FastText Embedding. Feature extraction algorithms, particularly BERT and RoBERTa, are scrutinized for their effectiveness in capturing nuanced text representations. Introducing a novel ensemble method, the research demonstrates its consistent superiority over individual models, achieving notable accuracy scores of 92.5% and 94.7% with BERT and RoBERTa, respectively. In an added layer of sophistication, hybrid feature selection involving LASSO and Recursive Feature Elimination (RFE) is incorporated post-feature extraction to refine and enhance the model’s performance. Through comprehensive comparative analyses, both pre and post hyperparameter tuning, the ensemble method asserts dominance over traditional machine learning classifiers, even surpassing Random Forest ensembles in accuracy and recall. Deep learning classifiers further validate the ensemble’s efficacy, achieving an impressive accuracy of 95.6% with RoBERTa. The study also reveals RoBERTa’s consistent slight advantage over BERT, aligning with established trends in natural language processing tasks. Beyond the research realm, these findings bear practical implications, as the proposed ensemble method, integrated with RoBERTa and enhanced through hybrid feature selection, emerges as a versatile and robust solution for accurate and nuanced emotion classification in diverse real-world scenarios, offering valuable insights for enhancing text classification performance.